Own intelligence. Kill the billable token.
Syzygy’s mission is to enable cost-efficient, outcomes-driven AI. We are a group of mathematics and physics researchers rebuilding AI deployment from the ground up as we think it should be: efficient, locally controlled, and owned rather than rented.
98%
of full-precision performance preserved by Mach-1 Small
160 tok/s
on devices with as little as 16 GB of unified memory
~1/10
the cost of the original deployment
<2 bits
per parameter: models and inference built for it
The problem
Renting intelligence is misaligned by design
Syzygy’s mission is to enable cost-efficient, outcomes-driven AI, and to kill the billable token. Today, inference providers of both open and closed-source models charge primarily for usage, whether per GPU hour or million tokens, rather than outcomes. Given the nondeterminism of token consumption per unit of work and the enormous energy, hardware, and labor costs of building and maintaining datacenters, this tradeoff is reasonable but painfully less than ideal. AI-native services mitigate this, aiming to more efficiently convert compute into outcomes by designing specialized products for specific industries. But that does not eliminate the upstream misalignment with compute providers, which is fundamental to the current model of “renting intelligence”.
We are already seeing the first signs in software engineering. Companies are spending extraordinary amounts on tokens approaching the cost of human labor itself, forcing them to cap token consumption per engineer. We predict this same tension will spread into any industry that begins to depend on AI.
Our thesis
Intelligence you own, like an operating system
Syzygy believes that cost-efficient, outcomes-driven AI becomes structurally feasible when users are able to own intelligence as easily as an Operating System. Once intelligence is owned rather than rented, we predict incentives will shift from metering units of reasoning to completing as much useful work as possible.
This is possible if the hardware and energy costs of AI begin to approach the cost of a modern personal computer. We see the bottleneck as four related problems:
capability / byte
Capability per byte of memory
capability / op
Capability per unit of computation
ops / watt
Computation per watt
work / dollar
Useful work per dollar of total system cost
These problems span the entire stack: model architecture, numerical representation, inference software, memory systems, and chip design. Syzygy is beginning with two foundational technologies:
models
Sub-2-bit language models
engines
Inference engines, and eventually chips, designed specifically for sub-2-bit computation
What we're building
Starting with sub-2-bit models and inference software
Most modern models and accelerators were built around relatively high-precision arithmetic. But inference does not require every parameter to be stored and processed at that precision. If model capability can be preserved below two bits, the cost of storing parameters, moving them through memory, and computing with them can fall dramatically.
We are starting with sub-2-bit models and inference software. Our proprietary compression algorithm and inference engine bring the capabilities of large language models to small, locally controlled devices. Our first model, Mach-1 Small, preserves 98% of the measured performance of its full-precision source model while supporting large contexts on devices with as little as 16 GB of unified memory. It runs at up to 160 tokens per second and approximately one-tenth the cost of the original deployment.
In the coming months, we plan to release Mach-1 XS, a smaller 1-bit model designed to run on mobile devices, and Mach-1 Medium, which brings the capabilities of a 120-billion-parameter model to laptops with 36 GB of unified memory.
Over the longer term, we plan to build affordable integrated systems capable of running large models entirely within enterprise environments.
The future
The future of AI will be heterogeneous
The future of AI will be heterogeneous. Some intelligence will run in hyperscale data centers. Some will run inside companies. Some will run on laptops, phones, robots, vehicles, and machines that have not yet been built.
Not every model will run locally, and not every workload should. The important question is whether people and companies have a genuine choice between renting intelligence and owning it.
We could not be more excited to launch Syzygy.
We are a group of mathematics and physics researchers rebuilding AI deployment from the ground up as we think it should be: efficient, locally controlled, and owned rather than rented. If that mission resonates with you, we invite you to join us.