Mach-1 Small benchmarks
Methodology, result tables, and limitations for published Mach-1 Small quality and local-performance claims.
Published result
In a hosted-GPU quality evaluation, the unweighted mean of Mach-1 Small's
twelve benchmark-level score-retention ratios was 95.0% versus its own BF16
teacher. Mach-1 Small (Mach-1-Ternary-Additive-35B) is a 35B-A3B MoE
compressed to about 1.7 bits per weight, in which every weight matmul is
add/subtract-only over integer codes; trained scale surfaces sit over frozen
integer codes via on-policy distillation. Its teacher is Qwen3.6-35B-A3B (BF16).
Retention is a compression-fidelity measure, not an absolute cross-model score:
retention = 100 × student score ÷ that student's own BF16 teacher scoreA value above 100 on one benchmark is possible because evaluation scores have variance.
The comparison
All three arms have figures for all twelve benchmarks, so there is one mean per model over one common set:
| Model | Mean retention (12 benchmarks) |
|---|---|
| Mach-1 Small | 95.0% |
| Ternary Bonsai 27B | 93.6% |
| Gemma 4 Q2_K_XL | 85.6% |
Gemma sits well back because it retains only about 67% on AIME25 and AIME26 — the selective collapse on sustained reasoning that the Bonsai whitepaper itself describes. On the ten benchmarks excluding AIME the same comparison is 94.2% / 93.9% / 89.3%, so almost the entire Gemma gap comes from those two rows.
Mach-1 Small and Ternary Bonsai are measured end to end on one harness. Every row in both columns is our own run, including AIME and tau2 on the fixed external simulator, so those two compare benchmark for benchmark with nothing borrowed. The Gemma column is external — see comparability.
The Gemma values are recomputed from Table 14 of
PrismML's Bonsai 27B whitepaper,
pinned at repository commit f904ea2ae3bb48e664de3ef36f55b553211b0c3c and
SHA-256 06451897df438a42d0f067a019db4e18969aa37aded7781a219f3b1fba54608b. The
ten non-AIME Gemma rows reproduce this arm's previously published student and
teacher scores exactly, which is what makes extending the same arm to AIME from
the same table sound.
The aggregate values and protocol summary are available as a machine-readable result summary. This is not a sample-level reproducibility package: exact artifact hashes, run IDs, sample counts, prompt hashes, and raw outputs are not yet public.
Protocol
| Item | Setting |
|---|---|
| Evaluation harness | PrismML App. B (prismml-app-b) |
| Evaluator | EvalScope 1.9.1 |
| Serving | vLLM OpenAI-compatible API |
| Hardware | Modal H200 × 2, tensor parallelism 2 |
| Mode | Thinking mode |
| Sampling | top_p=0.95, top_k=20 |
| Temperature | 1.0 for Mach/Qwen; 0.7 for Bonsai to match its published protocol |
| Arms run here | Mach-1 Small and Ternary Bonsai (the Gemma column is external) |
| Token tiers | PrismML tiers — 16,384 short; 20,480 medium; 30,000 long; 81,920 extended |
| AIME | Mean of 8 samples |
| Instruction metrics | IFEval prompt-strict, mean of 5 independent runs; IFBench prompt-loose |
| tau2 | Fixed external user-simulator (Qwen3.6-35B BF16, greedy), single pass |
This is a hosted-GPU quality evaluation. It does not measure Apple Silicon latency or throughput.
Per-benchmark retention
Ordered by Mach-1 Small's retention. Mach-1 Small and Ternary Bonsai were run here; the Gemma column is PrismML's published measurement throughout. See comparability below.
| Benchmark | Mach-1 Small / 35B teacher | Bonsai / 27B teacher | Gemma Q2_K_XL / Gemma teacher |
|---|---|---|---|
| AIME26 | 99.5% | 92.7% | 67.7% |
| MATH-500 | 99.4% | 98.2% | 95.6% |
| AIME25 | 99.1% | 91.7% | 67.2% |
| GSM8K | 98.4% | 100.2% | 97.3% |
| MBPP+ | 98.3% | 98.4% | 92.2% |
| HumanEval+ | 97.4% | 98.7% | 94.1% |
| MMLU-Redux | 96.2% | 94.0% | 96.9% |
| IFEval | 94.0% | 89.8% | 95.5% |
| MuSR | 92.7% | 91.6% | 91.1% |
| BFCL-v3 | 92.0% | 98.9% | 95.7% |
| tau2 | 90.0% | 91.2% | 73.1% |
| IFBench | 83.2% | 77.7% | 61.3% |
| Mean, 10 benchmarks excluding AIME | 94.2% | 93.9% | 89.3% |
| Mean, all 12 benchmarks | 95.0% | 93.6% | 85.6% |
Every Mach-1 Small and Ternary Bonsai figure above is our own measurement. Every Gemma figure is PrismML's published measurement — we did not run that model.
Underlying Mach-1 Small scores and teacher scores:
| Benchmark | Score | Teacher | Retention |
|---|---|---|---|
| AIME26 | 89.58 | 90.00 | 99.5% |
| MATH-500 | 98.00 | 98.60 | 99.4% |
| AIME25 | 87.50 | 88.33 | 99.1% |
| GSM8K | 94.69 | 96.21 | 98.4% |
| MBPP+ | 94.44 | 96.03 | 98.3% |
| HumanEval+ | 92.68 | 95.12 | 97.4% |
| MMLU-Redux | 89.18 | 92.68 | 96.2% |
| IFEval | 83.75 | 89.05 | 94.0% |
| MuSR | 61.77 | 66.66 | 92.7% |
| BFCL-v3 | 68.97 | 74.98 | 92.0% |
| tau2 | 71.58 | 79.51 | 90.0% |
| IFBench | 54.08 | 64.97 | 83.2% |
Comparability of the Bonsai and Gemma columns
Ternary Bonsai: measured here, all twelve rows. Nothing in that column is carried from the whitepaper any more. AIME25, AIME26, and tau2 on the fixed external simulator were all run on the same harness as Mach-1 Small, so the two columns are directly comparable benchmark for benchmark.
Where the whitepaper also publishes a figure for the same model, it runs higher than our measurement of it:
| our run | whitepaper | |
|---|---|---|
| AIME25 | 83.33 / 90.84 → 91.7% | 90.84 / 93.29 → 97.4% |
| AIME26 | 86.67 / 93.48 → 92.7% | 87.50 / 93.33 → 93.8% |
Some of that is the tau2 and AIME setups differing, and some is ordinary harness variance. The practical consequence is one-directional: published Ternary Bonsai numbers taken from the whitepaper are not comparable to this table, and mixing the two would flatter that column.
Gemma: entirely PrismML's published measurement. We did not run that model at all — every Gemma figure in this document is recomputed from Table 14. Its ten non-AIME rows reproduce the student and teacher scores we previously published for that arm exactly, which is what confirms the whole column has one source. Its tau2 in particular uses the whitepaper's own simulator, not the fixed external one behind the Mach and Bonsai figures.
One protocol difference remains inside our own run: IFEval is a mean of 5 independent runs for Mach-1 Small and a single run for Ternary Bonsai.
What the chart does and does not show
- Each compressed model is compared with its own teacher. The distance between models is not an absolute capability gap.
- The spread between Mach-1 Small and Ternary Bonsai is 1.4 points over twelve benchmarks but 0.3 over the ten excluding AIME. Almost all of it comes from AIME, where Mach-1 Small retains ~99% and Ternary Bonsai ~92%.
- The 95.0% vs 85.6% gap against Gemma is real but rests heavily on the two AIME rows, where Gemma retains about 67%. On the ten benchmarks excluding AIME the same gap is 94.2% vs 89.3%.
- Mach-1 Small is a 35B-A3B mixture-of-experts model; Ternary Bonsai is a dense 27B model. This is not a matched-architecture comparison.
- The Gemma values come from the Bonsai whitepaper, not the Mach/Bonsai same-harness run. They are ratios recomputed from the shared benchmark rows and are included as an external reference.
- The coding tasks used process-local execution rather than PrismML's sandbox. Same-harness ratios reduce those biases between the Mach and Bonsai arms, but do not eliminate all harness effects.
- Results can vary with prompts, sampling, evaluator versions, runtime, and hardware. A benchmark result is evidence for this configuration, not a guarantee for every task.
July 2026 local performance
The local figures were measured on Apple Silicon in a separate July 2026 run. Time to answer is prefill plus decode wall-clock time; lower is better.
| Prompt length | Mach-1 Small | Bonsai 27B | Gemma 4 Q2 |
|---|---|---|---|
| 128 tokens | 3.7 s | 6.1 s | 11.5 s |
| 2,048 tokens | 5.5 s | 9.8 s | 16.0 s |
| 8,192 tokens | 12.9 s | 22.3 s | 36.8 s |
"Intelligence per second" is an internal composite metric: each model's published retention mean multiplied by decode throughput from the local run. "Intelligence density per second" divides that value by measured model memory from the local run.
These are rescaled to the twelve-benchmark retention means above. Throughput and memory come from the July 2026 local run and did not change, so each value moved by exactly the ratio of new mean to old — which is why only the Gemma column moves: its mean fell from 89.3% to 85.6% once AIME was included, while the other two moved by a fraction of a point.
| Metric | Mach-1 Small | Bonsai 27B | Gemma 4 Q2 |
|---|---|---|---|
| Intelligence per second | 15.6 pts/s | 9.2 pts/s | 4.7 pts/s |
| Intelligence density per second | 1.97 pts/GB/s | 1.27 pts/GB/s | 0.40 pts/GB/s |
These measurements are indicative, not a guarantee for another machine or workload. Hardware, runtime, model artifacts, prompt and output length, cache state, thermal state, and repetition policy can change the result.