Joint engineering effort compresses weeks of low-level kernel optimization into a rapid iteration cycle, using autonomous tooling to map Qwen3 onto d-Matrix's SRAM-based Corsair architecture
SAN FRANCISCO--(BUSINESS WIRE)--#AI--Infinity (Infinity Artificial Intelligence Institute), an early-stage AI infrastructure research company building the software layer that makes any AI chip inference-ready, today announced a new case study that showcases tools for its autonomous research and tool building agent, Ignition, which generates, test and optimize the low-level compute kernels, compilers, profilers, debuggers and SDKs that determine how efficiently a chip runs AI models. Developed for a design partnership with d-Matrix for its SRAM-based inference accelerator Corsair, Infinity’s new AI product drastically reduces the time chip companies need before chips are ready for mass market adoption. Ignition is a concrete example of ongoing recursive self-improvement (RSI), as an AI system that builds and autonomously researches the training and inference layers for the next generation of AI systems.


New AI chips are frequently held back not by their hardware capabilities, but by the absence of a mature software stack, precisely what NVIDIA has spent the past two decades building around CUDA. That gap is what typically keeps promising accelerators out of production, and inference now accounts for a growing majority of AI compute spending industry-wide. Infinity’s new tools autonomously iterate on hardware representation and kernel design without requiring a large team of specialized kernel engineers and years of work for each new chip.
“Recursive self-improvement just delivered a scientific breakthrough that will upend the competitive landscape for chips. In a matter of weeks, we built a large part of an alternative to CUDA, which NVIDIA took 20 years to perfect. The tooling we developed for Infinity's Ignition in the process will speed up turnaround times for future design partnerships,” said Jeremy Nixon, founder and CEO of Infinity.
How the Result Was Achieved.
The d-Matrix Corsair chip is a memory-centric accelerator that keeps compute tightly integrated with on-chip SRAM, avoiding the memory-bandwidth bottlenecks that constrain GPU-based inference. Getting a model like Qwen3 to run efficiently on this architecture required mapping weights, activations and cache across a hierarchy of chiplets, gangs, slices, cores and SRAM banks — a memory-packing problem distinct from GPU optimization.
Infinity's engineering approach combined four techniques to distribute the workload:
- Pipelining: across the card's two packages, splitting model layers to minimize cross-package data transfer.
- Tensor parallelism: sharding dense compute (Q/K/V, output projection, and MLP layers) across 16 hardware gangs.
- Head-parallel attention: assigning each of Qwen3's attention heads to its own gang with no cross-gang communication required.
- Batch parallelism: isolating each sequence's KV cache and state to dedicated slices.
Given that Corsair's SRAM allows weights to stay resident across both prefill and decode phases, Infinity and d-Matrix built explicit lifetime tracking to keep expensive parameters in place while aggressively recycling transient activation memory, expanding usable capacity without repeatedly moving weights.
The optimization loop was supported by three purpose-built tools:
- Hardware probing and profiling, producing an empirical performance model of real instruction costs and memory behavior on Corsair.
- MemoryScope, Infinity's compiler, which searches memory placements and parallelism configurations for capacity, lifecycle and locality, paired with a Graph and Memory Sanitizer that statically verifies memory correctness before code runs on hardware.
- Agentic debugging, which drives breakpoints and memory inspection across the model's load-prefill-decode phases to localize numerical errors.
"As inference scales globally, customers need heterogeneous infrastructure where GPUs and purpose-built accelerators work seamlessly together. Getting there requires the ability to enable models faster on rack-scale hardware,” said Sid Sheth, founder and CEO of d-Matrix. “Working with Infinity, we were able to have models running on production-ready Corsair hardware in days, which means customers can deploy truly heterogeneous compute faster. That's a breakthrough for the entire ecosystem."
INFINITY.INC FREQUENTLY ASKED QUESTIONS:
What does Infinity do?
Infinity builds the software layer that enables any AI chip to run inference workloads. Its autonomous AI agent, Ignition, automatically generates and optimizes the inference software stack for new silicon in days, a process that has historically taken engineering teams months or years.
What is Ignition and how does it work?
Ignition is Infinity's AI research agent that autonomously writes, tests and optimizes the low-level compute kernels that determine how efficiently a chip runs AI models. It operates without human kernel engineers and continuously improves performance through a real-world feedback loop.
Why can't new AI chips compete with NVIDIA out of the box?
NVIDIA's CUDA ecosystem represents two decades of inference software development that rival chipmakers cannot quickly replicate. Most new chips fail not because of weak hardware, but because they lack the software stack to run AI models efficiently. Infinity's platform solves this by automating the software development entirely.
How fast can Infinity bring a new chip to production-ready AI inference?
In a published case study with chip partner d-Matrix, Infinity reached 92% of a new chip's theoretical peak performance within 10 hours of first hardware access and had three frontier models running end-to-end within 10 days. That same process has historically taken engineering teams months or years.
Which chip companies is Infinity working with?
Infinity's first live design partnership is with d-Matrix, whose Corsair chip now runs production-quality inference via the Infinity d-Matrix Cloud. Infinity is currently in active partnership discussions with other major chip companies.
Who founded Infinity.inc?
Infinity.inc was founded by Jeremy Nixon, a former Google Brain researcher, co-founder of the AGI Houses, a San Francisco and Hillsborough-based AI community and hacker network that has launched hundreds of startups and projects and a widely recognized voice on the future of AI, with his work and commentary covered by The New York Times and Forbes.
Is Infinity.inc hiring?
Hiring new talent is one of Infinity’s key goals after raising its $15M seed funding round. The company plans to expand its engineering team. Current open roles include research engineers working on hardware enablement and global inference library engineers maintaining Infinity’s Infy library.
About Infinity
Infinity.inc (Infinity Artificial Intelligence Institute) is an AI infrastructure research company building the software layer that makes any chip competitive for AI inference. Its AI research agent, Ignition, automatically generates, tests and optimizes the low-level compute kernels that determine how efficiently a chip runs AI models, compressing inference software development from years into days. Infinity's live design partnership with d-Matrix powers the Infinity d-Matrix Cloud, and the company is currently in active partnership discussions with other major chip companies.
Founded in August 2025 and headquartered in San Francisco, Infinity is a privately held company backed by Touring Capital and prominent angel investors, including executives at major chip companies, researchers from OpenAI and Anthropic, and Infinity CEO and founder, Jeremy Nixon. Follow Infinity.Inc on X and LinkedIn and Jeremy Nixon on X and LinkedIn, or learn more at https://infinity.inc.
Contacts
Media Contacts:
Mindy M. Hull
Mercury Global Partners for Infinity
+1 415 889 9977
press@infinity.inc
Michael Held-Hernandez
Mercury Global Partners for Infinity
+1 480 306 1154
press@infinity.inc





