About us · our thesis

Own intelligence. Kill the billable token.

Syzygy’s mission is to enable cost-efficient, outcomes-driven AI. We are a group of mathematics and physics researchers rebuilding AI deployment from the ground up as we think it should be: efficient, locally controlled, and owned rather than rented.

The problem

Renting intelligence is misaligned by design

Syzygy’s mission is to enable cost-efficient, outcomes-driven AI, and to kill the billable token. Today, inference providers of both open and closed-source models charge primarily for usage, whether per GPU hour or million tokens, rather than outcomes. Given the nondeterminism of token consumption per unit of work and the enormous energy, hardware, and labor costs of building and maintaining datacenters, this tradeoff is reasonable but painfully less than ideal. AI-native services mitigate this, aiming to more efficiently convert compute into outcomes by designing specialized products for specific industries. But that does not eliminate the upstream misalignment with compute providers, which is fundamental to the current model of “renting intelligence”.

We are already seeing the first signs in software engineering. Companies are spending extraordinary amounts on tokens approaching the cost of human labor itself, forcing them to cap token consumption per engineer. We predict this same tension will spread into any industry that begins to depend on AI.

Our thesis

Intelligence you own, like an operating system

Syzygy believes that cost-efficient, outcomes-driven AI becomes structurally feasible when users are able to own intelligence as easily as an Operating System. Once intelligence is owned rather than rented, we predict incentives will shift from metering units of reasoning to completing as much useful work as possible.

This is possible if the hardware and energy costs of AI begin to approach the cost of a modern personal computer. We see the bottleneck as four related problems:

capability / byte

Capability per byte of memory

capability / op

Capability per unit of computation

ops / watt

Computation per watt

work / dollar

Useful work per dollar of total system cost

These problems span the entire stack: model architecture, numerical representation, inference software, memory systems, and chip design. Syzygy is beginning with two foundational technologies:

models

Sub-2-bit language models

engines

Inference engines, and eventually chips, designed specifically for sub-2-bit computation

What we're building

Starting with sub-2-bit models and inference software

Most modern models and accelerators were built around relatively high-precision arithmetic. But inference does not require every parameter to be stored and processed at that precision. If model capability can be preserved below two bits, the cost of storing parameters, moving them through memory, and computing with them can fall dramatically.

We are starting with sub-2-bit models and inference software. Our proprietary compression algorithm and inference engine bring the capabilities of large language models to small, locally controlled devices. Our first model, Mach-1 Small, retains 96.3% of its BF16 teacher’s capability across twelve benchmarks and runs on supported Apple Silicon Macs with 16 GB or more of unified memory.

In the coming months, we plan to release Mach-1 XS, a smaller 1-bit model designed for mobile devices, and Mach-1 Medium, a higher-capability model designed for higher-memory laptops.

Over the longer term, we plan to build affordable integrated systems capable of running large models entirely within enterprise environments.

Mach-1 Smallavailable now
Its mean of twelve benchmark-level score-retention ratios was 96.3% versus its own BF16 teacher in a same-harness evaluation; supported on 16 GB and larger Apple Silicon Macs.
Mach-1 XScoming months
A smaller 1-bit model designed to run on mobile devices.
Mach-1 Mediumcoming months
A higher-capability local model designed for higher-memory laptops.
Integrated systemslonger term
Affordable integrated systems capable of running large models entirely within enterprise environments.

The future

The future of AI will be heterogeneous

The future of AI will be heterogeneous. Some intelligence will run in hyperscale data centers. Some will run inside companies. Some will run on laptops, phones, robots, vehicles, and machines that have not yet been built.

Not every model will run locally, and not every workload should. The important question is whether people and companies have a genuine choice between renting intelligence and owning it.

We could not be more excited to launch Syzygy.

We are a group of mathematics and physics researchers rebuilding AI deployment from the ground up as we think it should be: efficient, locally controlled, and owned rather than rented. If that mission resonates with you, we invite you to join us.