Fine-Tuning Financial Judgment Models: Bridgewater and Thinking Machines in Production

In July 2026, Bridgewater AIA Labs and Thinking Machines Lab published results that should change how every fintech CTO thinks about AI strategy. They fine-tuned Qwen3-235B — an open-weight MoE model — to achieve 84.7% accuracy on financial document triage, outperforming GPT-5.5, Claude Opus 4.8, and Gemini 3.1 Pro on the same benchmark. (See the CISPO loss paper on arXiv for the underlying training technique.) The headline number is not the accuracy — it is the economics. Their fine-tuned model delivered superior results at 13.8x lower inference cost than GPT-5.5. For financial institutions processing millions of documents per day, that difference moves from “interesting” to “existential” very quickly. ...

July 16, 2026 · 3 min · jnas