← Back to Jopopedia

Changelog

current v15 SFT July 2026

Golden Q&A benchmark (597 questions, gold context): EM 83.6% · F1 91.6% · shape 90.8%. End-to-end (live retrieval): reading top-3 passages lifted EM 24.1% → 30.8%; multi-turn follow-ups 11% → 37% via query rewriting. Retriever unchanged (v14r10).


v14r10 retriever July 2026

End-to-end benchmark (543 questions, live retrieval, no gold context): EM ~26% (parity with prior stack) · F1 +3pp · image EM 14% → ~35%. Aggregate EM is retrieval-bound — the model extracts at ~85% when handed the right passage; finding it at rank 1 among 48.7M passages is the open problem, and this release is the foundation for attacking it (cross-encoder v3 on filtered hard negatives, smarter confidence gating).


v14r4 + v13r2-ret June 2026

Golden Q&A benchmark (480 questions, gold context): EM 81.2% · F1 88.7% · shape-adherence 91.0% (v14r4 DPO step 200, deployed 2026-06-14). EM is intentionally a touch below v14's 83.8% — once answers are full sentences, exact-match punishes valid phrasing variety, so F1 + shape are the truer measure (both flat-or-up). Retriever unchanged (v13r2-ret); the v14 retriever retrain is the next lever.


v14 SFT + v13r2-ret June 2026

Golden Q&A benchmark (454 questions, gold context, corrected eval): EM 85.7% · F1 90.8% (v14 DPO step 500, deployed 2026-06-12; image bucket 77.5%, image-refusal bucket 100%, idk 91.7%). End-to-end via the unchanged v13r2-ret retriever: EM 23.8% · F1 41.7% (pre-DPO reading) — the attribution diagnostic shows rank-1 image extraction improved 23→38% while the retriever still fails to surface a usable infobox for 67% of image queries. The v14 retriever retrain is the next lever.


v13r13 + v13r2-ret June 2026

Golden Q&A benchmark (446 questions, gold context): EM 82.5% · F1 88.8% (v13r13b DPO step 500, deployed 2026-06-06). End-to-end via deployed v13r2-ret retriever (446 q, no gold context): EM 26.0% · F1 42.3% — flat vs prior v13 deploy and well below the gold-context ceiling. Image bucket is at 0% on n=40 (gold ceiling 65–67%); under investigation as the dominant drag.


v13r10 June 2026

Golden Q&A benchmark (407 questions, gold context): EM 79.1% (v13r6 deployed: 66.1%, +13.0pp) · F1 85.0% (v13r6 deployed: 75.2%, +9.8pp)


v13 May 2026

Golden Q&A benchmark (78 questions, gold context): EM 47.4% (v12+DPO deployed: 42.3%, +5.1pp) · F1 67.7% (v12+DPO: 66.3%, +1.4pp) · biggest gains in when (+29 EM), dimension (+25 EM), what_did (+25 EM), founder (+50 EM)


v12 May 2026

SFT benchmark (gold-context, 500q): SQuAD EM 17.6% (v11: 17.4%) · Internal v11 EM 60.8% (+16.0pp over v12 SFT baseline) · Internal v12 EM 64.4% (+14.6pp over baseline) · Retrieval passage_match@3 76.0% (v11 stack: 47.6%) · infobox/disambig retrieval went from 0% to 85%+ since v11 didn't have those passage types in its index at all


v11 May 2026

v11 capability benchmark: 96% overall (v10: 40%) · SFT eval/loss: 0.0973 (v10: 0.1107, −12%) · Bi-encoder recall@10: 96% · Cross-encoder hard-val acc: 91.9%


v10 April 2026

Internal EM: 58.4% · Internal F1: 85.0% · Golden QA EM: 43.4% (53 questions) · SQuAD EM: 16.6%


v9r3 April 2026

Val loss: 0.1048 · Internal EM: 40.8% · Internal F1: 66.8%


v8 April 2026

SQuAD EM: 19.0% (new high) · Internal EM: 47.4% (new high) · Internal F1: 73.7% (new high, +3.3 over v6r2)


v7 April 2026


v6r2 April 2026

SQuAD EM: 18.8% (+92% over v6) · Internal F1: 70.4% (best ever)


v6 April 2026

SQuAD EM: 9.8% · Internal EM: 46.4% · Internal F1: 68.2%


v5 March 2026

SQuAD EM: 10.8% · Internal EM: 45.6% (gold context)


v4 March 2026

SQuAD EM: 7.8% (+39x over v3)


v3 March 2026

Data quality was the biggest win: +42.8% EM over v2


v2 September 2025


v1 August 2025