AI Inference Silicon
- Exponential Industry
- Technology Area
- 22 Milestones
Dedicated application-specific processors and accelerators engineered for high-throughput, low-latency AI model inference in data centers and edge deployments. Architectures overcome the memory wall by utilizing massive on-chip SRAM, high-bandwidth memory (HBM3e), and specialized matrix processing engines. Commercial implementations include Etched Sohu ASICs, Groq Language Processing Units (LPUs), Tenstorrent Wormhole, and Cerebras CS-3 systems.
Exponential Industry · Hidden Factory
AI Inference Silicon
NVIDIA Groq 3 LPX Enters Full Production; Nebius Deploys in Token Factory
NVIDIA announced full production of the rack-scale Groq 3 LPX system (256 LPUs, 128GB SRAM, 640TB/s bandwidth), with Nebius deploying clusters into the Nebius Token Factory for low-latency agentic decoding achieving 3,400 tok/s on Gemma 4 31B /NVIDIA Newsroom/ /Nebius/.
OpenAI Unpacks Jalapeño Custom LLM Inference ASIC at Hot Chips 2026
OpenAI presented architecture and benchmark results for Jalapeño at Hot Chips 2026, detailing its reticle-sized 64-slice inference ASIC with 216 GB HBM4, spatial programming model, and 9-month RTL-to-tapeout development cycle with Broadcom /OpenAI/ /Tom's Hardware/.
Ornn OCPI still prices A100, H100, H200, B200, and RTX 5090
Ornn Compute Price Index hourly API last_updated 2026-08-22T14:02:53Z: A100 SXM4 $1.06, H100 SXM $2.85, H200 $4.91, B200 $6.73, RTX 5090 $0.53 per GPU-hour. Settled daily index 2026-08-21T20:00:00Z puts RTX 5090 at $0.52. Free-tier docs list exactly these five GPUs. Snapshot of active rental-market financial impact — not a claim that any of them is still the frontier pretrain SKU /Ornn/.
Etched delivers first inference rack to Jane Street
Etched PR Aug 18, 2026: first rack shipped last month; Jane Street said it tested the chip, is pleased with early results, and has its own rack running in its data center. Treated as the announcement day; ship month is July 2026 without a calendar day /GlobeNewswire/.
Meta MTIA 300 presented at ISCA 2026
MTIA Team (Meta Platforms) paper on MTIA 300, Meta's first training chip with built-in NIC chiplets and collective offloading engines, listed for ISCA 2026 (Raleigh, main program Jun 29–Jul 1). start_date is the ISCA 2026 main-program start, not a sourced individual talk slot /Meta Platforms / ISCA 2026/.
Vector Core fully disaggregated inference demonstration
Xeon 6 orchestration, SambaNova SN40 decode, NVIDIA Blackwell prefill from Vector Core LA data center; Together.ai first commercial customer.
US Patent Office Publishes OpenAI Patent US20260147536A1 on Hardware Alignment Accelerators
The USPTO published OpenAI's patent US20260147536A1 ('Alignment in Hardware Accelerators') describing data alignment mechanisms and mantissa scaling in compute-in-memory (CIM) macro arrays for high-efficiency neural accelerators /US Patent and Trademark Office/ /X/.
Moore Threads $1.1B Shanghai Star Board IPO
Moore Threads listed on the Shanghai Star Board raising $1.1 billion with a 500% first-day rally and 4,086x retail oversubscription /Bloomberg/.
NVIDIA documents Unsloth QLoRA fine-tune on GeForce RTX 5090
NVIDIA Technical Blog (Oct 23, 2025) reports Unsloth QLoRA fine-tuning on a GeForce RTX 5090 (32GB, Alpaca, Llama 3.1 8B-class table) and states a single Blackwell GPU can fine-tune models with as many as 40 billion parameters. This is the sourced 5090 training use-case — adapter fine-tune, not frontier pretrain. Explains why 5090 still clears a ~$0.52–0.53/hr Ornn spot without having trained Llama/GPT-class base models /NVIDIA Technical Blog/.
GeForce RTX 5090 available at $1,999
CES PR (Jan 6, 2025) stated GeForce RTX 5090 would be available on Jan 30 at $1,999 (3,352 AI TOPS). Shipping day for the consumer Blackwell SKU that later appears on Ornn OCPI as the cheapest free-tier GPU /NVIDIA Newsroom/.
NVIDIA GeForce RTX 5090 announced at CES 2025
CES 2025 (Jan 6): NVIDIA unveiled GeForce RTX 50 Series on Blackwell. RTX 5090 is 92 billion transistors and over 3,352 AI TOPS; NVIDIA said it would be available Jan 30 at $1,999. Consumer SKU — local inference and creator AI, not a data-center HBM training GPU. No sourced SOTA foundation model was pretrained on RTX 5090 /NVIDIA Newsroom/ /NVIDIA GeForce News/.
Meta Llama 3.1 405B trained on over 16 thousand H100 GPUs
Meta AI blog (Jul 23, 2024) introducing Llama 3.1, including 405B as the first frontier-level open-source Llama. Meta said it pushed training to over 16 thousand H100 GPUs on 15+ trillion tokens, making 405B the first Llama trained at that scale. Primary sourced breakthrough-model evidence for H100's Ornn price — the Hopper fleet that trained the 2024 open frontier model still rents in 2026 /Meta AI/.
NVIDIA Blackwell platform and GB200 NVL72 announcement (GTC 2024)
GTC 2024 (Mar 18) launch of the NVIDIA Blackwell platform: Blackwell GPU architecture, GB200 Grace Blackwell Superchip, liquid-cooled GB200 NVL72 rack-scale system (72 Blackwell GPUs + 36 Grace CPUs on fifth-gen NVLink), HGX B200, sixth-generation air-cooled DGX B200 (8 Blackwell GPUs, two 5th Gen Intel Xeon; up to 144 PFLOPS FP4), and broad cloud/OEM adoption commitments. Foundational rackscale-AI milestone that established the NVL72 rack product line later continued by GB300 and Vera Rubin NVL72; DGX B200 is the traditional air-cooled 8-GPU SuperPOD/BasePOD building block, not NVL72 /NVIDIA Newsroom/.
NVIDIA H200 HGX announced (SC23)
SC23 (Nov 13, 2023): NVIDIA announced HGX H200, the first GPU with HBM3e (141GB at 4.8 TB/s; nearly double A100 capacity and 2.4x bandwidth). NVIDIA said H200 would nearly double Llama 2 70B inference vs H100. Availability from 2Q 2024 as a compatible HGX H100 successor. Hopper memory-refresh branch, not a new architecture /NVIDIA Newsroom/.
Meta Llama 2 paper reports pretraining on NVIDIA A100-80GB
arXiv 2307.09288 (Jul 18, 2023): Meta released Llama 2 (7B–70B). Pretraining used NVIDIA A100s on RSC and internal production clusters; a cumulative 3.3M GPU hours on A100-80GB (TDP 400W or 350W). This is the sourced breakthrough-model evidence for A100's remaining financial impact — not that A100 still trains 2026 frontier runs, but that the open-weight Llama 2 generation ran on the Ampere fleet /arXiv/.
NVIDIA H100 Hopper GPU in full production
GTC Fall 2022 (Sep 20): NVIDIA announced H100 in full production, with partners planning an October rollout of Hopper-based products and services. DGX H100 (eight H100, 32 petaflops FP8) available to order. This is when Hopper became a rentable/shipped fleet, not only a GTC slide /NVIDIA Newsroom/.
NVIDIA H100 Transformer Engine blog / Hopper TE introduction coverage
NVIDIA Blog (Dave Salvator) details Transformer Engine on Hopper H100: mixed FP8/FP16 Tensor Core training/inference for transformers, MoE 395B and Megatron 530B speedups vs prior gen. Updated Aug 8, 2023 to note Ada Lovelace incorporation of Transformer Engine.
NVIDIA Hopper architecture and H100 announced (GTC 2022)
GTC 2022 (Mar 22): NVIDIA announced Hopper, succeeding Ampere, and the first Hopper GPU H100 (80 billion transistors, Transformer Engine, HBM3, fourth-gen NVLink). Jensen Huang called H100 the engine of the world's AI infrastructure as data centers become AI factories. Availability stated as starting in the third quarter of 2022 /NVIDIA Newsroom/.
NVIDIA A100 Ampere GPU in full production (GTC 2020)
GTC 2020 (May 14): NVIDIA announced the first Ampere-architecture GPU, A100, in full production and shipping. More than 54 billion transistors on 7nm; unifies AI training and inference with up to 20x vs prior gen; MIG up to seven instances; third-gen NVLink. DGX A100 (eight A100s) announced the same day. Trunk of the modern NVIDIA LLM-training GPU line that Ornn still prices in 2026 /NVIDIA Newsroom/.
Andy Grove becomes Intel CEO
Grove became president 1979, CEO 1987, chairman 1997–2005. Under his leadership Intel produced the 386 and Pentium; annual revenue rose from $1.9B to more than $26B (span not dated in the PR) /Intel/.
Intel operations in Oregon begin
Innovating and Investing in Oregon Since 1974 /Intel/.
Intel founded; Grove first hire
Noyce and Moore left Fairchild to found Intel in 1968; Grove was their first hire /Intel/.
Frequently Asked Questions
What is AI Inference Silicon and what industrial engineering problems does it address?
Dedicated application-specific processors and accelerators engineered for high-throughput, low-latency AI model inference in data centers and edge deployments. Architectures overcome the memory wall by utilizing massive on-chip SRAM, high-bandwidth memory (HBM3e), and specialized matrix processing engines. Commercial implementations include Etched Sohu ASICs, Groq Language Processing Units (LPUs), Tenstorrent Wormhole, and Cerebras CS-3 systems.
What core technologies and manufacturing architectures comprise AI Inference Silicon?
Key technical architectures and manufacturing innovations include:
- NVIDIA Blackwell Platform: Full-stack accelerated computing platform (GPU architecture, NVLink, networking, software) for trillion-parameter generative AI training and inference. Succeeds Hopper; named for mathematician David Harold Blackwell. Consumer GeForce RTX 50 Series (RTX 5090 CES 2025-01-06) is the same architecture family.
- NVIDIA Hopper Architecture: NVIDIA data-center GPU architecture announced GTC 2022-03-22, succeeding Ampere. First SKU H100 (Transformer Engine, HBM3); H200 (2023-11-13) is the HBM3e memory refresh. Still the training fleet for Llama 3.1-class models; Ornn still prices H100 SXM and H200 as free-tier OCPI GPUs in Aug 2026.
- NVIDIA A100 Tensor Core GPU: Ampere-architecture data-center GPU in full production and shipping 2020-05-14 (>54 billion transistors, 7nm, MIG, third-gen NVLink, TF32). Meta pretrained Llama 2 on A100-80GB clusters (3.3M GPU hours; arXiv 2307.09288, 2023-07-18). Ornn OCPI A100 SXM4 $1.06/hr as of 2026-08-22T14:02:53Z — cheapest data-center SKU on the free-tier index because a large depreciated Ampere fleet still runs fine-tune, inference, and smaller training jobs after Hopper/Blackwell took frontier pretrain.
- NVIDIA H100 Tensor Core GPU: Hopper-architecture data-center GPU announced GTC 2022-03-22 (80 billion transistors, Transformer Engine, first HBM3). Full production 2022-09-20. Meta trained Llama 3.1 405B on over 16 thousand H100 GPUs (Meta blog, 2024-07-23). Ornn OCPI H100 SXM $2.85/hr as of 2026-08-22T14:02:53Z. SemiAnalysis InferenceX (NVIDIA DGX B200 page, as of Apr 2026) cites Hopper inference at about $0.09/MTok vs Blackwell $0.02/MTok on GPT-OSS-120B — remaining financial impact is the installed training/inference fleet, not leading token TCO.
- NVIDIA GeForce RTX 5090: Consumer Blackwell GeForce GPU announced CES 2025-01-06 (92 billion transistors, 3,352 AI TOPS, 32GB GDDR7, 21,760 CUDA cores, 575W TGP) and available 2025-01-30 at $1,999. NVIDIA markets it for local LLM inference (LM Studio, AnythingLLM) and NVIDIA's Oct 23, 2025 Unsloth blog documents QLoRA fine-tune of Llama 3.1 8B-class models on a single 5090, with a claim that a single Blackwell GPU can fine-tune up to 40B-parameter models. No sourced evidence that any frontier/SOTA foundation model was pretrained on RTX 5090 — 32GB GDDR7 is not an HBM training SKU. Ornn OCPI $0.53/hr hourly (2026-08-22T14:02:53Z) and $0.52/hr settled daily (2026-08-21) — cheapest of the five free-tier GPUs, reflecting consumer/local-AI and small-job rental rather than cluster pretrain.
- Etched frontier inference cluster: Rack-scale inference cluster using Low Voltage Inference (LVI) for compute density at the same power and Cluster Scale Memory (CSM), a shared memory pool across the cluster. First customer Jane Street; >$1B in contracts claimed.
Which companies and industrial facilities lead deployment in AI Inference Silicon?
Leading industrial manufacturers, hyperscalers, and engineering operators include:
- NVIDIA Corporation: Meta Llama 3.1 405B trained on over 16 thousand H100 GPUs (Jul 23, 2024) — To enable training runs at this scale and achieve the results we have in a reasonable amount of time, we significantly optimized our full training stack and pushed our model training to over 16 thousand H100 GPUs, making the 405B the first Llama model trained at this scale.
- Meta: Meta MTIA 300 presented at ISCA 2026 — We present MTIA 300, Meta's first AI training chip optimized for Deep Learning Recommendation Models (DLRMs).
- Intel: Vector Core fully disaggregated inference demonstration (Jun 2, 2026) — Xeon 6 orchestration, SambaNova SN40 decode, NVIDIA Blackwell prefill from Vector Core LA data center; Together.ai first commercial customer.
- OpenAI: US Patent Office Publishes OpenAI Patent US20260147536A1 on Hardware Alignment Accelerators (May 28, 2026) — Techniques for data alignment in hardware accelerators, including aligning mantissa bits and exponent scaling within compute-in-memory (CIM) macro arrays for high-efficiency neural network inference.
- Broadcom: OpenAI Unpacks Jalapeño Custom LLM Inference ASIC at Hot Chips 2026 — Jalapeño is OpenAI's custom inference ASIC co-designed with Broadcom in a nine-month RTL-to-tapeout development cycle. Optimized for LLM inference efficiency and throughput per watt to overcome data center power constraints.
- Etched: Etched delivers first inference rack to Jane Street (Aug 18, 2026) — Etched shipped its first rack last month to Jane Street, and the quantitative trading firm is actively deploying the technology into its workloads.
What are key commercial projects and milestone achievements in AI Inference Silicon?
Major industrial breakthroughs and commercial milestones include:
- NVIDIA A100 Ampere GPU in full production (GTC 2020): NVIDIA today announced that the first GPU based on the NVIDIA Ampere architecture, the NVIDIA A100, is in full production and shipping to customers worldwide.
- NVIDIA Hopper architecture and H100 announced (GTC 2022): Named for Grace Hopper, a pioneering U.S. computer scientist, the new architecture succeeds the NVIDIA Ampere architecture, launched two years ago.
- NVIDIA Blackwell platform and GB200 NVL72 announcement (GTC 2024): NVIDIA today announced that the NVIDIA Blackwell platform has arrived — enabling organizations everywhere to build and run real-time generative AI on trillion-parameter large language models at up to 25x less cost and energy consumption than its predecessor.
- Meta Llama 3.1 405B trained on over 16 thousand H100 GPUs (Jul 23, 2024): To enable training runs at this scale and achieve the results we have in a reasonable amount of time, we significantly optimized our full training stack and pushed our model training to over 16 thousand H100 GPUs, making the 405B the first Llama model trained at this scale.
- US Patent Office Publishes OpenAI Patent US20260147536A1 on Hardware Alignment Accelerators (May 28, 2026): Techniques for data alignment in hardware accelerators, including aligning mantissa bits and exponent scaling within compute-in-memory (CIM) macro arrays for high-efficiency neural network inference.
- OpenAI Unpacks Jalapeño Custom LLM Inference ASIC at Hot Chips 2026: Jalapeño is OpenAI's custom inference ASIC co-designed with Broadcom in a nine-month RTL-to-tapeout development cycle. Optimized for LLM inference efficiency and throughput per watt to overcome data center power constraints.
Related Technology Ontologies
Physical AI & Embodied Robotics
Artificial intelligence systems and autonomous machines that operate directly in and interact with the physical world, bridging bits to atoms. Spanning foundation models (Vision-Language-Action and World Foundation Models), physics-based simulation with synthetic data (Omniverse, Isaac Sim) to safely bridge Sim2Real, and embedded runtime computers (Jetson Thor, DRIVE AGX) executing closed-loop perception-action loops across humanoids, AMRs, adaptive manipulators, autonomous mobility, and smart spaces.
Data Center Power & Behind-the-Meter Generation
Generation and storage sited at or next to AI/cloud campuses so the load does not wait on a transmission queue. Common stacks pair gas turbines or engines with on-site BESS; some designs add behind-the-meter solar. Hyperscalers and developers announced multi-hundred-megawatt gas+BESS campuses in 2025–2026 to serve rack-scale AI halls.
Rackscale AI Accelerators
Fully co-designed rack- and POD-scale AI systems that combine accelerators, host CPUs, scale-up fabric (NVLink-class domains), and scale-out networking in one thermal/mechanical package. Lineage runs from multi-GPU HGX/DGX boxes through Grace Hopper NVL32 domains to liquid-cooled 72-GPU racks (GB200, GB300, Vera Rubin) and competing racks such as AMD Helios. NVIDIA, AMD, and OEM rack builders (Dell, Supermicro) ship the current generation for AI-factory halls.