SK hynix Says HBM Alone Won’t Carry AI Inference Much Further

SK hynix showed off HBF, PIM, and SALT-KV at AI Infra Summit 2026: a NAND tier borrowing HBM's stacking, aimed at inference, not training.

SK hynix used AI Infra Summit 2026 in Santa Clara to argue that the memory hierarchy behind AI inference needs a layer that doesn't exist yet. Its answer is HBF, High Bandwidth Flash: a NAND product built with the same through-silicon via stacking HBM uses, sitting somewhere between HBM's bandwidth and an SSD's capacity.

The reasoning traces back to where AI workloads have shifted. Training rewards raw bandwidth, which is what HBM sells. Inference, especially long-context and agentic workloads, needs to keep far more state available than HBM can hold at a sane cost, while an SSD is too slow to feed a GPU directly. HBF is pitched squarely at that gap, and SK hynix says it drew more visitor interest than anything else on its stand.

The SK hynix stand at AI Infra Summit 2026, with separate SALT-KV and PIM demonstration stations.
Two of the three technologies on one stand: the SALT-KV station on the left, the PIM station beside it. SK hynix says the HBF model drew the most visitor interest of anything it showed. Image: SK hynix.

Three Products Aimed at the Same Bottleneck

Alongside HBF, the company showed AiM, its processing-in-memory chip, and the AiMX accelerator card built from it, running a live large language model demo. Processing-in-memory puts computation inside the memory itself, so less data has to move between processor and memory: a direct attack on the memory wall rather than an attempt to widen the pipe.

The third piece was SALT-KV, short for Semantic-Aware Lifecycle Tiering for KV Cache. A language model's key-value cache grows with conversation length and task count, and it grows fast. SALT-KV splits that cache into context-based segments, weighs each segment's likely reuse against what it costs to store, and places it in HBM, DRAM, or SSD accordingly. It's a software answer to the same problem HBF attacks in hardware.

Vice President Lim Eui-cheol, who heads Solution AT, laid out the case for all three in a session presentation. "As AI workloads grow more diverse, ranging from ultra-low-latency agent services to long-context services where cost efficiency is critical, the conventional GPU-HBM architecture alone is increasingly unable to meet new requirements," he said. SK hynix is preparing "new solutions that span both hardware and software: PIM, HBF, and SALT-KV."

What This Is, and What It Is Not

This was a trade-show stand and a conference session, not a product launch. SK hynix named no capacities, no bandwidth figures, no sampling dates, and no customers for HBF, and gave no pricing for anything shown. The claim that PIM and HBF are "excellent alternatives that can move beyond existing limits" is the company's own assessment of its own roadmap, made on its own stage, with no independent testing behind it.

What's worth taking from this is the direction, not the specifics. The largest memory maker in the HBM business is telling its own customers that HBM alone won't carry inference, and is spending stand space on the tier below it. The event itself drew roughly 6,000 attendees by SK hynix's count, double the previous year, a sign of how much of the industry is now arguing about the same bottleneck.

Sources