Subquadratic launches with $29M to build AI with 12M-token context windows
What's the deal? Subquadratic, a startup building a new type of generative AI model, has emerged from stealth with $29M in seed funding. Its model, SubQ, can process up to 12 million tokens at once — roughly 120 books — compared with the industry standard of 128,000 tokens or up to 1 million for frontier models like Claude and Gemini.
The company was co-founded by chief executive Justin Dangel and chief technology officer Alexander Whedon. SubQ uses a proprietary transformer architecture built on sparse attention, which avoids comparing every token to every other token — the costly default approach in most large language models.
Why now? Traditional transformers use dense attention, where doubling the input roughly quadruples the compute required. That "quadratic" scaling problem has kept context windows small relative to the data volumes many enterprises need to process. Subquadratic says its linear scaling approach means doubling the input only doubles the compute.
The company claims SubQ is more than 50 times faster and 50 times cheaper than leading frontier models at 1 million tokens, while maintaining higher accuracy. At its full 12 million-token window, it says compute requirements drop by nearly 1,000 times compared with other frontier models.
What could go wrong? Sparse attention architectures involve trade-offs. Skipping token-to-token comparisons can sacrifice nuance, and real-world performance on complex reasoning tasks may differ from benchmarks. The company also enters a crowded field where deep-pocketed incumbents — OpenAI, Google, and Anthropic — are racing to expand their own context windows.
A $29M seed round, while substantial, is modest by AI standards. Scaling infrastructure and winning enterprise customers against well-funded rivals will test the startup's resources.
The signal: Context window size is becoming a key competitive axis in AI. Enterprises want models that can ingest entire codebases, legal archives, or financial datasets in a single pass — without ballooning costs. Subquadratic's bet is that the architecture itself, not just scale, is the bottleneck worth solving.
If sparse attention delivers on its promise, it could reshape how companies think about retrieval-augmented generation and other workarounds built to compensate for short context windows. The broader trend: AI infrastructure is shifting from "bigger models" to "smarter architectures."
Read more: siliconangle.com