Subquadratic says sparse attention cut LLM cost to $8, not $2,600
Independent Appen tests back speed and coding performance, but wide availability is still the missing piece.

Miami startup Subquadratic, via CEO Justin Dangel and CTO Alex Whedon, claims its SubQ model breaks a decades-old transformer bottleneck using sparse attention. If its cost and speed hold up in broader use, procurement and product teams could rethink how they budget and deploy frontier coding and document-heavy LLM workloads.
Subquadratic came out of stealth last month with a claim that immediately divided AI twitter: it says it cracked the mathematical bottleneck that makes today’s large language models expensive and power-hungry. Now the company is putting numbers behind that story, including a cost comparison that is hard to ignore. Subquadratic says it costs $8 to run its SubQ model through Nvidia’s RULER 128 benchmark, versus $2600 to run Anthropic’s LLM Opus 4.6 through the same test.
That same push for receipts is showing up in independent evaluation too. SubQ reportedly scored 89.7% on LiveCodeBench, a benchmark focused on competitive coding problems from real contests, and Appen’s testing also includes a speed result: in a straight-up speed test, SubQ was 56 times faster than models using FlashAttention. The headline stake for decision-makers is simple: if SubQ can deliver frontier-level coding performance with dramatically lower compute, then “how much does this call cost?” stops being a minor line item and becomes a board-level architecture choice.
To understand why this could matter, you have to zoom out to how most LLMs work today. The core mechanism is a transformer running dense attention. Dense attention, in plain terms, multiplies token-to-token relationships across the entire input. The result is quadratic compute growth as context gets longer. The source illustrates the computation problem with a simple intuition: if you double the number of words, you roughly quadruple the computations, which is why long-context workloads can turn into power and budget nightmares.
Subquadratic’s pitch is that it ditches dense attention and uses sparse attention, selecting only some relationships between tokens instead of multiplying everything with everything. The company says sparse attention is not new in spirit. The challenge, historically, has been that previous sparse-attention approaches often use fixed patterns, like always comparing the first word to the fifth. Subquadratic argues that those fixed patterns are too limiting because language is too dynamic. In its approach, the selection is calculated on the fly and differs for each piece of text, which Whedon describes as the “secret sauce.” SubQ is positioned as the first sparse-attention LLM that can rival mainstream dense-attention models in performance.
And this is where the skepticism initially landed. Subquadratic’s first announcement reportedly had thin details, with people unconvinced by self-published test scores and not enough independent verification. The model was also not widely available for others to try. That combination is basically how you trigger a “prove it” moment in an industry where benchmarks are competitive and incentives are intense.
A month later, Subquadratic started to bring the receipts by publishing more information about SubQ, including results from Appen, an evaluation firm that was brought in to test the model. Subquadratic cofounder and chief technology officer Alex Whedon said in the source that releasing third-party benchmarks alongside the initial announcement would have preempted skepticism and that the company is taking time to ensure future results are fully verified before putting them out. Appen’s Jeanine Sinanan-Singh, director of generative AI research, is quoted describing the significance: she says the results validated SubQ’s architecture and called the speed and inefficiency angle potentially game changing, while also noting that self-claims become less credible when outcomes are “shocking.”
For executives, the non-obvious angle is what these tests imply about workload shape, not just model identity. SubQ is not positioned as a universal replacement for top models across the board. Instead, it’s framed as offering “huge increases in speed at a fraction of the typical cost” for certain tasks. That matters because many real businesses do not run one benchmark at a time. They run specific workflows: coding assistance, document analysis, and retrieval-style tasks where context length and latency determine whether the product feels magical or slow. The source also claims SubQ can process up to 12 times as much text at once than most other models, enabling data-heavy tasks like analyzing hundreds of documents or entire code bases.
There are still gaps to watch. Subquadratic’s cost claims are harder to verify because SubQ is not yet widely available. The company’s RULER 128 comparison is specific, but broad procurement decisions usually require repeated trials under real constraints, with your own data, your own tokenization patterns, and your own integration stack. Also, the source notes that for context windows, SubQ’s window is up to 12 million tokens, while most top models today are around one million tokens. Longer context is valuable, but it only pays off if retrieval and reasoning behave reliably when inputs get huge.
The second-order implication for peers is that “efficiency” could become a competitive axis again, not just a performance race. If SubQ truly delivers frontier-level coding performance while cutting compute cost and energy use, then teams building coding tools, developer platforms, and enterprise copilots may start budgeting around sparse-attention style architectures, not only the biggest parameter counts. Subquadratic’s CEO Justin Dangel says the company hopes it is kicking off “a new age of efficiency,” and that it does not think anyone will be building on transformers in a few years. Whether that timeline is right or not, the underlying decision problem is already here: when cost drops from thousands of dollars per benchmark run to single digits, your unit economics change, and product roadmaps tend to follow.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Business

Anthropic’s Levant Alpöge cracks the Jacobian conjecture after 87 years
A Harvard valedictorian used Claude to hit a 1939 breakthrough, but the missing “why” is the real problem.

Uber buys Delivery Hero for nearly $15B, vaulting to top food delivery outside China
The deal doubles Uber's dual-services footprint and pushes a ride-and-eats bundling play into 50 more markets.

Epic and Google drop settlement bid, forcing rival Android app stores by July 22
Google told the court it is ready to carry third-party app stores starting Wednesday, July 22.

