All research

AI · Annual review · 2026-08-30

The Bottleneck Moved

Memory became an allocation calendar, optics received real money, neoclouds became capital structures, and the scarce product moved from a GPU to useful output per active megawatt.

, Founder and publisher, AI Bottlenecks

TL;DR

The scoreboard, with a warning label

The current AI Bottlenecks universe contains 120 report-date names across 16 themes. Usable one-year history existed for 119. The model equal-weights themes daily, then names inside each theme.

This universe was built after the measurement period began. Projecting it backward creates look-ahead, survivorship, rebalancing and global-calendar effects. The figures show where today's thesis would have found sensitivity. They are not a live portfolio return.

| Theme | One-year return | Drawdown from peak | | --- | ---: | ---: | | Memory Supercycle | +852.9% | -24.7% | | InP & Substrates | +487.4% | -26.8% | | Photonics / CPO | +325.6% | -24.5% | | HBM / Packaging | +252.5% | -22.2% | | Lithography & Fab Tools | +178.9% | -31.0% | | Custom Silicon | +174.9% | -24.2% | | Networking / Retimers | +119.3% | -27.9% | | Power & Grid | +85.8% | -34.0% | | AI Cloud / Neoclouds | +81.9% | -22.4% | | S&P 500 | +19.4% | n/a |

The ordering matters more than the headline return. The four strongest groups sat below or beside the accelerator, where a small amount of industrial capacity governed a much larger amount of compute revenue.

!Return and drawdown across the physical AI stack

The year in six moves

1. Memory became an allocation calendar

The year began like a normal memory recovery. Prices improved, inventories normalized and margins came back. Then the disclosures stopped looking cyclical.

Micron's fiscal 2025 revenue rose to $37.38 billion from $25.11 billion, while gross margin rose to 39.8% from 22.4%. Its cloud-memory unit produced $4.54 billion in the fourth quarter, more than triple the prior year, at a 59% gross margin. The physical mechanism mattered more: Micron said an HBM bit used roughly three times the wafer area of a DDR5 bit, with the ratio increasing in future generations. Micron FY2025 | Micron prepared remarks

That ratio explains why HBM supply can expand while ordinary DRAM remains tight. HBM competes for front-end wafers before it reaches thinning, TSVs, bonding and qualification. The PC buyer is competing with the wafer-equivalent appetite of every AI bit.

SK hynix said demand for its entire 2026 DRAM and NAND production was secured, then moved HBM4 into mass shipment. Samsung still described memory supply as insufficient in July despite weaker consumer end markets. The early signal had moved from a spot quote to how far ahead a customer had to commit. SK hynix Q2 2026 | Samsung Q2 2026

2. The package started editing the roadmap

HBM4 added a logic base die, taller stacks, harder thermals, tighter alignment and platform-specific qualification. TSMC's roadmap made the surrounding geometry explicit: CoWoS had reached a 5.5-reticle package in production, while the 2028 target reached 14 reticles, ten compute dies and 20 HBM stacks. The package was becoming the machine that decided how many expensive known-good parts could operate together. TSMC

That is why the tool orders were cleaner proof than another demand forecast:

  1. Onto Innovation disclosed double-digit orders for HBM4 inspection systems.
  2. Camtek announced more than $105 million of metrology orders from a leading HBM manufacturer and a tier-one OSAT.
  3. FormFactor said HBM and CPO drove sequential growth.
  4. Besi's first-half orders rose 116.5%, with photonics, hybrid bonding and AI among the drivers.

Onto | Camtek | FormFactor | Besi

Once a package contains several expensive known-good components, test is not overhead. It is the act of avoiding a late-stage scrap event that destroys the value already inside the stack.

3. Optics received money

NVIDIA launched its CPO switch roadmap in March 2025. The change this year was that the roadmap turned into transaction consideration, equity, deposits and factory access.

Meanwhile, Lumentum's fiscal 2026 revenue rose 83.2% to $3.014 billion and Fabrinet's rose 36% to $4.64 billion. Current pluggables and future CPO were feeding the same factories.

The clean read is not that CPO killed pluggables. Faster pluggables carried current revenue while CPO reserved tomorrow's capacity. Both increased the value of the laser, substrate, package and test chain.

4. NVIDIA started manufacturing bankability

The popular recap is that NVIDIA went shopping. The transaction record is stranger and more important.

SchedMD was a conventional acquisition. Groq was not. Groq licensed inference technology to NVIDIA while key executives and engineers joined NVIDIA; Groq remained independent. NVIDIA recorded $13 billion paid at close, another $4 billion due within a year, $14.4 billion of goodwill and $2.5 billion of developed technology. Calling it a normal license understates the substance. Calling it an acquisition is inaccurate. Groq | NVIDIA accounting note

At 26 July, NVIDIA disclosed $279 billion of supply and capacity commitments, $99 billion of equity investments, $25 billion of future equity commitments, $36 billion of AI-cloud service commitments and $20 billion of third-party leases not yet commenced. It also disclosed selected guarantees, including a phased $105 billion maximum cap for SB Energy after quarter-end. These are not all cash paid or present losses. They show how much balance sheet now sits around the silicon. NVIDIA Q2 FY2027 10-Q

!An AI factory runs on six industrial clocks

NVIDIA is closing the handoffs where a sold accelerator can fail to become revenue-bearing capacity: memory, optics, cloud distribution, scheduling, inference and project finance. The moat is stretching from benchmark leadership into the ability to make the whole deployment bankable.

5. Neoclouds became capital structures

CoreWeave, Nebius, IREN and Applied Digital no longer fit inside “GPU rental.” Each combines some portion of cloud operator, capacity developer and financing vehicle.

| Company | Operating receipt | Capacity or finance receipt | | --- | --- | --- | | CoreWeave | Q2 revenue $2.575B, adjusted EBITDA $1.510B, net loss $626M | 1.5 GW active; 3.7 GW contracted; quarterly interest expense $640M | | Nebius | Q2 revenue $582.3M, +454% | Year-end contracted-power target of 5 GW; some cohorts 50% to 60% pre-funded | | IREN | FY2026 AI Cloud revenue $128.8M | GPU finance around 6% for Microsoft and 9% for weaker-credit cohorts | | Applied Digital | Roughly 100 MW operating at 31 May | $2.35B of 9.25% notes; 1,410 MW contracted across five campuses |

CoreWeave is the clearest proof that demand and equity pain can coexist. Its quarterly interest expense exceeded its net loss. Nebius showed the cleaner contract-backed version, including a $17.4 billion Microsoft agreement and Meta agreements with potential value around $27 billion. The contracts improve financeability. They do not energize a data hall.

A better ladder is: announced GW, secured power, financed capacity, contracted compute, active MW, utilized compute, cash-generating output. A company can move several rungs in a press release and zero rungs in revenue.

6. Inference changed the denominator

Google said Gemini serving unit costs fell 78% during 2025 while direct API traffic rose above 10 billion tokens per minute. Lower unit cost and much higher total demand happened together. Alphabet Q4 2025 call

> Revenue per MW = tokens per second per MW x realized revenue per token x utilization

Then subtract power, networking, operations, depreciation and financing.

Prefill and decode also began to separate. Long prompts and reasoning traces make prefill compute-heavy. Interactive generation values low latency and predictable token speed. Rubin can handle broad throughput, Groq LPX can target fast generation, Dynamo can route the work and Slurm can govern the cluster below it.

The valuable layer is increasingly the one deciding where the workload runs.

!The bottleneck kept moving one layer away from the obvious trade

Sivers and the difference between a thesis and proof

The Sivers thesis had a real technical core. Ayar Labs used Sivers DFB laser arrays. WIN Semiconductors was its outsourced high-volume manufacturing partner. O-Net, Enablence, POET, Jabil and GlobalFoundries represented various routes into laser sources, light engines, pluggable modules and reference designs.

The broader social claim was not proven. Reviewed primary sources did not establish that Sivers supplied every major optical-I/O winner, or confirmed supply to Lightmatter and Celestial AI.

The financial gap matters. Sivers reported 2025 sales of SEK306.6 million, up 40%, while adjusted EBITDA remained negative SEK50.3 million. In Q1 2026, its opportunity pipeline reached $799 million, but sales fell 22% year over year and operating cash flow was negative SEK49.2 million. Annual report | Q1 2026

That does not make the thesis fake. It defines the proof still required: production orders, photonics revenue, margin conversion, positive cash flow and clean reporting. The stock moved before the company economics did.

Five calls for the next 12 to 24 months

1. HBM keeps ordinary DRAM tighter than consumer units imply

HBM's wafer-area penalty and long-term agreements keep pulling capacity away from ordinary products.

Watch: wafer allocation, HBM prepayments and supplier contract duration. Wrong if: ordinary DRAM inventories rise while HBM shipments keep expanding without affecting pricing.

2. Test intensity grows faster than shipped HBM bits

Larger packages and more known-good dies increase the value of inspection, probe and bonding control.

Watch: HBM4 orders at Camtek, Onto, FormFactor and Besi. Wrong if: two quarters of HBM volume growth arrive with broad tool-order contraction.

3. Pluggables remain the larger optical revenue pool through 2027

CPO ramps, but current transceiver demand funds the same industrial base first.

Watch: 1.6T shipments, laser utilization, CPO production revenue and service data. Wrong if: CPO becomes material revenue faster than pluggables grow.

4. Active MW replaces contracted GW in neocloud valuation

Investors start pricing connected clusters, customer acceptance and utilization against interest and refresh obligations.

Watch: active MW, acceptance dates, revenue per MW and financing spread. Wrong if: contracted power keeps receiving full credit while commissioning slips.

5. NVIDIA becomes a recurring compute offtaker

Cloud commitments and guarantees increasingly support third-party capacity that also expands NVIDIA's installed base.

Watch: cloud-service commitments, guarantees, partner utilization and concentration. Wrong if: commitments shrink without new instruments replacing them.

The monitoring board

| Signal | Strong evidence | Warning | | --- | --- | --- | | HBM wafer allocation | Long agreements and tight ordinary DRAM | Inventory rises across products | | HBM tool orders | Inspection and probe convert with volume | Two quarters of broad contraction | | Optical prepayments | Deposits become shipments and revenue | Capacity grows without utilization | | CPO production | Measurable revenue, yield and field data | Demo count rises, revenue does not | | Active MW | Accepted clusters with stable utilization | Contracted GW rises while active MW slips | | Financing spread | Lower cost and diversified customers | Higher rates, guarantees and concentration | | Output per MW | Realized throughput and economics improve | Benchmarks rise without fleet returns |

The read

The best AI infrastructure trade was not simply more compute. It was the discovery that each unit of compute required more industrial coordination than the market had modeled.

Memory became a wafer-allocation problem. Photonics became a laser-capacity and package-reliability problem. Neoclouds became a credit and construction problem. Inference became a routing and utilization problem. NVIDIA's moat stretched into each handoff because any handoff could stop a sold chip from becoming productive capacity.

The scarce product is no longer a GPU in a warehouse. It is a powered, connected, financed and accepted system that stays busy doing useful work.

> Useful output per active MW, after depreciation and financing

Everything before that is potential. That moving boundary is what AI Bottlenecks is built to follow.

Methodology

Market figures cover 28 August 2025 through 28 August 2026 and project the report-date universe backward. Prices are Yahoo adjusted closes converted to U.S. dollars. The model equal-weights themes, then names within them, rebalanced daily. It excludes fees and factor adjustment and is descriptive, not a live performance record.

Selected sources

  1. Micron FY2025 results
  2. SK hynix Q2 2026
  3. TSMC advanced-packaging roadmap
  4. NVIDIA and Coherent
  5. NVIDIA and Lumentum
  6. AXT and Lumentum
  7. Groq and NVIDIA
  8. NVIDIA Q2 FY2027 filing
  9. CoreWeave Q2 2026
  10. Nebius Q2 2026
  11. Broadcom Q2 FY2026
  12. Alphabet Q4 2025 earnings call