{"schemaVersion":2,"evidenceKind":"aggregate-result-summary","claim":"Mach-1 Small's unweighted mean of twelve benchmark-level score-retention ratios was 95.0% versus its own BF16 teacher, against 93.6% for Ternary Bonsai 27B and 85.6% for Gemma 4 Q2_K_XL on the same twelve.","evaluationCompletedAt":"2026-08-03","publishedAt":"2026-08-03","scope":{"qualityOnly":true,"speedBenchmark":false,"crossModelAbsoluteCapabilityComparison":false},"protocol":{"scope":"Mach-1 Small and Ternary Bonsai arms (the Gemma arm is external)","harness":"PrismML App. B (prismml-app-b)","evaluator":"EvalScope 1.9.1","server":"vLLM OpenAI-compatible API","hardware":"Modal H200 x2, tensor parallelism 2","thinkingMode":true,"decoding":{"topP":0.95,"topK":20,"machTemperature":1,"bonsaiTemperature":0.7},"tokenTiers":{"short":16384,"medium":20480,"long":30000,"extended":81920},"perBenchmark":{"aime":"Mean of 8 samples.","ifeval":"Prompt-strict; mean of 5 independent runs.","ifbench":"Prompt-loose.","tau2":"Fixed external user-simulator (Qwen3.6-35B BF16, greedy), single pass, identical for every model in the run."}},"aggregation":{"formula":"100 * student_score / its_own_teacher_score","mean":"unweighted arithmetic mean of the twelve retention percentages","excluded":[]},"arms":{"mach":{"id":"mach1_ternary_additive_35b","displayName":"Mach-1 Small","modelId":"Mach-1-Ternary-Additive-35B","description":"35B-A3B MoE compressed to ~1.7 bits/weight; every weight matmul is add/subtract-only over integer codes (multiplier-free); trained scale surfaces over frozen integer codes via on-policy distillation.","teacherId":"Qwen3.6-35B-A3B BF16"},"bonsai":{"id":"ternary_bonsai_27b","displayName":"Ternary Bonsai 27B","teacherId":"Qwen3.6-27B BF16","note":"All twelve rows are our own same-harness run, including AIME25, AIME26, and tau2 on the fixed external simulator. Nothing here is carried from the whitepaper."},"gemma":{"id":"gemma_4_31b_q2_k_xl","displayName":"Gemma 4 Q2_K_XL","teacherId":"Gemma 4 31B BF16","source":{"title":"Bonsai 27B whitepaper","repositoryCommit":"f904ea2ae3bb48e664de3ef36f55b553211b0c3c","sha256":"06451897df438a42d0f067a019db4e18969aa37aded7781a219f3b1fba54608b","url":"https://github.com/PrismML-Eng/Bonsai-demo/blob/f904ea2ae3bb48e664de3ef36f55b553211b0c3c/bonsai-27b-whitepaper.pdf","table":"Table 14, full per-benchmark results, thinking mode","note":"Ratios recomputed from Table 14 for all twelve benchmarks Mach measured, including AIME25 and AIME26. Verified: the ten non-AIME rows reproduce this arm's previously published student and teacher scores exactly."}}},"results":[{"benchmark":"AIME26","note":"Mach and Bonsai: our run, mean of 8 samples. Gemma from Table 14 of the whitepaper. The whitepaper reports 87.50 for Bonsai here against a 93.33 teacher (93.8%); our harness puts the same model at 86.67 against a 93.48 teacher.","mach":{"student":89.58,"teacher":90,"retentionPercent":99.5},"bonsai":{"student":86.67,"teacher":93.48,"retentionPercent":92.7},"gemma":{"student":59.8,"teacher":88.33,"retentionPercent":67.7}},{"benchmark":"MATH-500","mach":{"student":98,"teacher":98.6,"retentionPercent":99.4},"bonsai":{"student":96.4,"teacher":98.2,"retentionPercent":98.2},"gemma":{"student":95.2,"teacher":99.6,"retentionPercent":95.6}},{"benchmark":"AIME25","note":"Mach and Bonsai: our run, mean of 8 samples. Gemma from Table 14 of the whitepaper. The whitepaper reports 90.84 for Bonsai here against a 93.29 teacher (97.4%); our harness puts the same model at 83.33 against a 90.84 teacher — 5.7 points of retention lower.","mach":{"student":87.5,"teacher":88.33,"retentionPercent":99.1},"bonsai":{"student":83.33,"teacher":90.84,"retentionPercent":91.7},"gemma":{"student":58.2,"teacher":86.66,"retentionPercent":67.2}},{"benchmark":"GSM8K","mach":{"student":94.69,"teacher":96.21,"retentionPercent":98.4},"bonsai":{"student":97.04,"teacher":96.82,"retentionPercent":100.2},"gemma":{"student":94.9,"teacher":97.57,"retentionPercent":97.3}},{"benchmark":"MBPP+","mach":{"student":94.44,"teacher":96.03,"retentionPercent":98.3},"bonsai":{"student":96.03,"teacher":97.62,"retentionPercent":98.4},"gemma":{"student":77.8,"teacher":84.39,"retentionPercent":92.2}},{"benchmark":"HumanEval+","mach":{"student":92.68,"teacher":95.12,"retentionPercent":97.4},"bonsai":{"student":94.51,"teacher":95.73,"retentionPercent":98.7},"gemma":{"student":90.7,"teacher":96.34,"retentionPercent":94.1}},{"benchmark":"MMLU-Redux","mach":{"student":89.18,"teacher":92.68,"retentionPercent":96.2},"bonsai":{"student":87.82,"teacher":93.47,"retentionPercent":94},"gemma":{"student":90.7,"teacher":93.6,"retentionPercent":96.9}},{"benchmark":"IFEval","note":"Mach: prompt-strict, mean of 5 independent runs.","mach":{"student":83.75,"teacher":89.05,"retentionPercent":94},"bonsai":{"student":81.52,"teacher":90.74,"retentionPercent":89.8},"gemma":{"student":86.5,"teacher":90.57,"retentionPercent":95.5}},{"benchmark":"MuSR","mach":{"student":61.77,"teacher":66.66,"retentionPercent":92.7},"bonsai":{"student":65.21,"teacher":71.16,"retentionPercent":91.6},"gemma":{"student":64.7,"teacher":71.03,"retentionPercent":91.1}},{"benchmark":"BFCL-v3","mach":{"student":68.97,"teacher":74.98,"retentionPercent":92},"bonsai":{"student":74.6,"teacher":75.46,"retentionPercent":98.9},"gemma":{"student":71.5,"teacher":74.71,"retentionPercent":95.7}},{"benchmark":"tau2","note":"Mach and Bonsai: our run on the fixed external user-simulator (Qwen3.6-35B BF16, greedy), single pass. The Gemma figure is the whitepaper's own tau2 setup.","mach":{"student":71.58,"teacher":79.51,"retentionPercent":90},"bonsai":{"student":69.53,"teacher":76.26,"retentionPercent":91.2},"gemma":{"student":53.2,"teacher":72.8,"retentionPercent":73.1}},{"benchmark":"IFBench","note":"Prompt-loose.","mach":{"student":54.08,"teacher":64.97,"retentionPercent":83.2},"bonsai":{"student":52.88,"teacher":68.03,"retentionPercent":77.7},"gemma":{"student":48.6,"teacher":79.3,"retentionPercent":61.3}}],"means":{"benchmarkCount":12,"machRetentionPercent":95,"bonsaiRetentionPercent":93.6,"gemmaRetentionPercent":85.6},"publicEvidenceAvailability":{"aggregateScores":true,"protocolSummary":true,"exactArtifactRevisionsOrChecksums":false,"runIds":false,"sampleCountsAndPromptHashes":false,"rawSampleOutputs":false},"caveats":["Each student is divided by its own teacher; the aggregate is not an absolute capability score across model families.","The Mach and Bonsai columns are now one harness end to end, including tau2 on the fixed external simulator, so they compare benchmark for benchmark with no borrowed rows.","Where the whitepaper also reports a figure, it runs higher than our measurement of the same model: AIME25 97.4% against our 91.7%, AIME26 93.8% against our 92.7%. Published Bonsai numbers elsewhere may not be comparable to this table.","The Gemma column is sourced entirely from PrismML's published measurements — we did not run that model. Its tau2 in particular uses the whitepaper's own simulator rather than the fixed external one used for Mach and Bonsai.","The Bonsai and Gemma figures come from the earlier run, before the current per-benchmark protocol. The Mach tau2 teacher moved from 78.06 to 79.51 between runs, so the tau2 row in particular spans two user-simulator configurations.","Mach-1 Small is a 35B-A3B mixture-of-experts model while Ternary Bonsai is a dense 27B model.","The Gemma series is recomputed from PrismML's Bonsai 27B whitepaper and was not produced by the Mach/Bonsai same-harness run.","Temperature differs by arm to match the published Bonsai protocol.","Coding used process-local execution rather than PrismML's sandbox.","The public evidence is an aggregate result summary, not a sample-level reproducibility package.","This summary does not support Apple Silicon throughput, latency, energy, or cost claims."]}