← Back to Jopopedia
Changelog
current v15 SFT July 2026
- It can hold a conversation now. Say "Hi", ask who it is, thank it, say goodbye — it responds naturally instead of trying (and failing) to look the greeting up in Wikipedia. Ask "Who are you?" and it'll tell you it's Jopopedia, a 520M-parameter model trained from scratch on Wikipedia.
- New question types trained. This release taught the model four new skills from scratch — multi-turn follow-ups ("...and when was it founded?"), picking the right passage out of several, "how does X work?" explanations, and the conversational behaviors above. On a harder, broader benchmark it matched the previous version's overall accuracy while adding all of these (behavioral 0→88%, multi-turn 48→70%, multi-passage 63→75%), with answer style staying tight (91% on-target verbosity).
- Cleaner answers. Fixed a long-standing bug that stuck an "A" in front of names ("A Leonardo da Vinci…" → "Leonardo da Vinci…"), and image requests now understand bare phrasings like "Show me Einstein" (not just "Show me a photo of Einstein").
- Now reads multiple sources per answer (live). The site now shows the model the top few retrieved passages instead of just one, and it picks the answer-bearing one — so answers that used to fail because the best passage ranked #2 or #3 now land ("Who founded Microsoft?" → "Bill Gates and Paul Allen."). End-to-end exact-match jumped +6.7 points across the board.
- Follow-up questions work now. Ask "What is the capital of Japan?" then "What about Brazil?" and it answers Brasília — it carries the topic across turns. Behind the scenes the follow-up is rewritten into a full question ("What is the capital of Brazil?") before answering, which turned out to work far better than making the model juggle the conversation itself.
- Under the hood: a full data regeneration (~2M training examples with the new capabilities woven in), and the model trained noticeably longer than usual — it kept finding small improvements out to step 12,000, so we let it.
Golden Q&A benchmark (597 questions, gold context): EM 83.6% · F1 91.6% · shape 90.8%. End-to-end (live retrieval): reading top-3 passages lifted EM 24.1% → 30.8%; multi-turn follow-ups 11% → 37% via query rewriting. Retriever unchanged (v14r10).
v14r10 retriever July 2026
- New bi-encoder trained by distillation, not contrastive learning — the retriever's embedding model is now a student of the cross-encoder: instead of the binary in-batch InfoNCE objective (which saturates early and then drifts), it learns to reproduce the cross-encoder's graded relevance scores over hard candidate sets, plus one ANCE round (re-mining harder negatives from the improved encoder and re-distilling). First encoder to beat the deployed one on every retrieval diagnostic: recall@50 91.3 vs 90.1, recall@1 72.9 vs 72.4, MRR 79.3 vs 78.2.
- Image and multi-meaning retrieval finally work — the new encoder wins 6 of 8 query-type buckets, with the biggest gains exactly where the old retriever was weakest: image lookups ("show me a photo of the Eiffel Tower" now finds the Paris original's infobox, not a same-name film or the Texas replica) and bare ambiguous names. End-to-end image EM roughly doubled (14% → ~35%).
- Same-name routing fixes — "Who is Marie Curie?" retrieves the scientist (not the charity), "Who is Genghis Khan?" the emperor (not the 1950 film), "Who founded Microsoft?" the actual founding paragraph (previously not even retrieved).
- Fresh 48.7M-passage index — rebuilt on the v14r10 corpus (latest prose + infobox + Wikidata + disambiguation passages).
- Confidence gate recalibrated — the new encoder scores on a different cosine scale, so the "I couldn't find that" threshold moved 0.75 → 0.55. (Found the hard way: the old threshold silently refused half of all queries.)
- Cross-encoder kept — it was the distillation teacher, so it stays aligned with the student. Two full retrain attempts on harder mined negatives both improved validation accuracy yet lost end-to-end — a lesson in "the metric isn't the product."
- Also since the last entry: the GPT itself advanced v14r4 → v14r9 → v14r10 (eval-accuracy fixes let training run longer; golden EM 86.2 / F1 92.9 on gold context, deployed as raw SFT).
- July 4 follow-up — cross-encoder v3 + title boost (the big one). The re-ranker was retrained on hard negatives mined from the live index and filtered by the old model's own judgment (~8% of mined "negatives" actually answered the question and had to be removed — training on them was what sank two earlier retrain attempts). Stacked with a one-line fix: any passage whose exact title appears in your question now gets a decisive ranking boost, so "Show me a picture of Albert Einstein" finds Albert Einstein instead of Einstein (unit). Together: end-to-end EM 25.4% → 34.6%, hard questions 20% → 32.5%, image lookups 35.7% → 69% with nearly half of all image answers now resolving to a real photo (vs ~20% this morning, ~10% in June).
End-to-end benchmark (543 questions, live retrieval, no gold context): EM ~26% (parity with prior stack) · F1 +3pp · image EM 14% → ~35%. Aggregate EM is retrieval-bound — the model extracts at ~85% when handed the right passage; finding it at rank 1 among 48.7M passages is the open problem, and this release is the foundation for attacking it (cross-encoder v3 on filtered hard negatives, smarter confidence gating).
v14r4 + v13r2-ret June 2026
- Answers now read like sentences — factual lookups come back conversationally ("There are 7 continents.", "Pulp Fiction was directed by Quentin Tarantino.", "The capital of Brazil is Brasília.") instead of bare fragments. A settled per-question-type verbosity policy (terse facts → one restating sentence; "who/what is" → one rich sentence) is enforced in the training data, the gold answers, and a new shape-adherence metric. On-style rate jumped 57% → 91% of correct answers, with answer F1 unchanged (88.7%).
- Dates standardized everywhere — every date renders as "March 20, 2022" regardless of the source format ("20 March 1727", "1727-03-20", etc. all normalize). Applied at answer-generation time across all data sources and made format-agnostic in scoring.
- Data-quality + image fixes — list-join cleanup ("X, and and Y" → "X, and Y"), founder questions no longer answered with a bare month/year, and image Q&A now skips non-renderable (.tif) and malformed filenames so every image answer is a real, extractable URL. New
audit_extractability.py verifies every answer is recoverable from its context as a standard post-regen check.
- v14r4 SFT + DPO — full regen on the new policies; SFT with primary-answer weighting so the target shape is actually learned, then a gentle DPO pass (snapshot sweep picked step 200).
Golden Q&A benchmark (480 questions, gold context): EM 81.2% · F1 88.7% · shape-adherence 91.0% (v14r4 DPO step 200, deployed 2026-06-14). EM is intentionally a touch below v14's 83.8% — once answers are full sentences, exact-match punishes valid phrasing variety, so F1 + shape are the truer measure (both flat-or-up). Retriever unchanged (v13r2-ret); the v14 retriever retrain is the next lever.
v14 SFT + v13r2-ret June 2026
- Image questions stop hallucinating URLs — 2.2M new negative-context training pairs teach the model to refuse when the retrieved passage has no
image: field instead of fabricating a plausible-looking filename (79% of emitted image URLs were fake in the v13r13 end-to-end eval; on gold context that's now zero). Image question phrasings doubled 10→20 ("Can you show me a photo of X?", "Picture of X please", "X photo please") with per-entity sampling so all 20 appear corpus-wide at unchanged volume.
- Same-name disambiguation descriptors — media infobox passages now lead with a year/type/attribution fragment right after the title:
Heart (2006 film). 2006 film by Hanny Saputra. director ... Covers film/album/song/book/play/video game at 95–99%; gives both the SFT model and the next retriever discriminating tokens for the Heart-vs-Heartbreakers collision class. Film needed two tricks: director as attribution (release dates didn't survive infobox extraction) and year-from-title fallback.
- Bare ambiguous queries: 3× the training signal — disambig question templates expanded 4→12 ("What could X refer to?", "Tell me about X", "X — what is it?"), 3.15M pairs from 263K disambiguation pages, aimed at the bare_ambiguous retrieval bucket.
- Retriever training hygiene — the bi-encoder no longer trains on deliberately-wrong pairs (IDK swap-passages, new image negatives) as positives — a quiet ~1-2% noise bug present since v12. Excluded across all retriever-side consumers (training, cross-encoder data, recall diagnostics, index building).
- Golden eval grows refusal coverage — 8 new image-negative entries (image question + prose context → expected refusal), 454 total. These steer checkpoint selection during SFT, not just measurement.
- v14 SFT + DPO — same 25K-cap recipe as v13r13b on the rebalanced 2.3M-entry set; best.pt at step 6500. DPO pass (6 miners incl. a dedicated "It is X."-prefix miner, 1,824 pairs, gentle v12 recipe) deployed at step 500.
- Standalone golden eval bug found + fixed — the standalone evaluator assumed every golden context carried a "Context: " prefix; 125 of 454 didn't and reached the model as malformed prompts, understating every standalone score by 3–5pp (the in-training evaluator was right all along). The prompt builder now enforces the prefix by checking the string, not trusting the caller.
Golden Q&A benchmark (454 questions, gold context, corrected eval): EM 85.7% · F1 90.8% (v14 DPO step 500, deployed 2026-06-12; image bucket 77.5%, image-refusal bucket 100%, idk 91.7%). End-to-end via the unchanged v13r2-ret retriever: EM 23.8% · F1 41.7% (pre-DPO reading) — the attribution diagnostic shows rank-1 image extraction improved 23→38% while the retriever still fails to surface a usable infobox for 67% of image queries. The v14 retriever retrain is the next lever.
v13r13 + v13r2-ret June 2026
- v13r13 SFT — 15 generation-layer fixes — landmark/element definitions kept their defining clause ("liquid at standard conditions" for Mercury), birth/died life-span paren extractor ("(1879–1955)" emits both born AND died Q&A), Oxford-comma list parser ("Steve Jobs, Steve Wozniak, and Ronald Wayne" no longer drops Wayne), disambig de-dup, multi-clause "How does X work?" capture, "In what year/location" templates, infobox+Wikidata leads drop the echo-friendly "X is an entry of type Y" lead, 20 IDK refusal phrasings (up from 7).
- Wikidata qid→label cache — one-shot pass over the Wikidata bz2 dump resolves 1.38M wikibase-item labels, unlocking P36 (capital), P37 (language), P38 (currency), P61 (inventor/discoverer), P170 (painter). New templates: "What is the capital of X?", "What language is spoken in X?", "Who invented/discovered/painted X?".
- v13 retriever: rebuilt from scratch, format-aligned — new branched data pipeline just for the retriever (own rebalance, amplifies primary_topic + def_who_is + def_what_is, drops insufficient_passage). Passages use a leading-metadata layout (
Title, "Section", paragraph N: body) that anchors title at position 0 for the bi-encoder; the SFT model still wants v13r9 trailing, so the deploy converts at the retrieval→generation boundary. Single unified 47.9M-passage index over Wikipedia prose (41.7M) + synthetic infobox (5.5M) + Wikidata (404K) + disambig (263K).
- Bi-encoder: phase 2 hard-neg contrastive — init from the reusable MLM 250K-step checkpoint. Phase 1 in-batch only at bs=88, 100K steps; phase 2 with partial-corpus FAISS-mined hard negatives (1M-passage subsample against the phase-1 encoder, 300K queries, ~50 min) and 85/15 mix at bs=80, ran 20K steps. Recall@20 climbed 91.8→96.0% on a held-out 30K-passage val set; the encoder was still improving at the cap so v14 will extend further.
- Cross-encoder retrained — v13r2ret-fixed-20k, bs=48 with
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True, format-converted training data so the CE sees the same leading layout the bi-encoder serves. Val acc 98.5% / val loss 0.054 at step 20K.
Golden Q&A benchmark (446 questions, gold context): EM 82.5% · F1 88.8% (v13r13b DPO step 500, deployed 2026-06-06). End-to-end via deployed v13r2-ret retriever (446 q, no gold context): EM 26.0% · F1 42.3% — flat vs prior v13 deploy and well below the gold-context ceiling. Image bucket is at 0% on n=40 (gold ceiling 65–67%); under investigation as the dominant drag.
v13r10 June 2026
- Context format change (v13r9 SFT) — passage metadata moved from a leading
Title, section N, "Section", paragraph M: body prefix to a trailing body (from "Section" of Title, paragraph M) parenthetical. The old prefix was an attractor for the model's induction heads: it would emit the article title as the first answer token and then copy the rest of the metadata verbatim ("Albert Einstein, section 1, \"Introduction\", paragraph 1: ..."). Catastrophic at later checkpoints; gone entirely now.
- v13r10 DPO — three new miners (degenerate repetition, disambig list quality, wrong-number consistency) seeded ~570 preference pairs from v13r9 outputs against training golds. v12 gentle recipe (lr=1e-7, beta=0.05, 600 steps, dpo_ratio=0.05) with 95% SFT mixing. Snapshot sweep picked step 600 as best.
- Eval normalize() upgrade — subject-aware wrap-strip, prefix-substring tolerance, verb-form fold for "How does X work", Oxford-list connector drop. Catches valid wrapped/verbose model outputs that the old eval punished as form mismatches.
Golden Q&A benchmark (407 questions, gold context): EM 79.1% (v13r6 deployed: 66.1%, +13.0pp) · F1 85.0% (v13r6 deployed: 75.2%, +9.8pp)
v13 May 2026
- 22 new infobox types — sportsperson, nrhp, election, french_commune, organization, cricketer, river, artist, mountain, building, military unit, baseball biography, football club, nfl biography, uk place, writer, university, military conflict, basketball biography, radio station, ice hockey player, rugby biography, artwork. Coverage of in-scope infoboxes jumped from 47% (v12) to 81% (v13). 11.1M total infobox Q&A pairs (up from 4.7M)
- Image lookup — "Show me a photo of Lincoln" returns a Special:FilePath URL. 2.56M Wikipedia infoboxes have an image; ~1.28M have a caption. 10 question phrasings ("Show me a photo of X", "Picture of X please", "What does X look like?", etc.) train the model on the new pattern. Frontend rendering of embedded images coming next.
- Definitional question variants — "Tell me about X", "Describe X", "Explain X" now route to the same definitional answer as "Who is X?" / "What is X?". Fixes the v12 production failure where "Tell me about king k rule" matched the archaeology term "Tell"
- Parser-based infobox extraction —
extract_infoboxes.py rewritten on top of mwparserfromhell + 22-worker multiprocessing. Cleanup time dropped from ~hours to 52 minutes for 4.6M articles. Audit shows all straggler patterns <0.01% (vs v12's 5%+ for some classes)
- v13 SFT + DPO recipe — 25K SFT steps on the new rebalanced 67M-entry training set, then a 2000-step DPO with the v12-gentle recipe (lr=1e-7, beta=0.05, snapshot sweep). v13+DPO step 2000 deployed
Golden Q&A benchmark (78 questions, gold context): EM 47.4% (v12+DPO deployed: 42.3%, +5.1pp) · F1 67.7% (v12+DPO: 66.3%, +1.4pp) · biggest gains in when (+29 EM), dimension (+25 EM), what_did (+25 EM), founder (+50 EM)
v12 May 2026
- Infobox-driven Q&A pipeline — structured Wikipedia infoboxes (settlement, person, officeholder, company, film) now feed the training set with 4.66M precise Q&A pairs. "Who directed Inception?", "What is the population of Tokyo?", "When was Tesla founded?" land on clean key-value facts instead of regex-extracted prose
- Multi-mode DPO post-training — four failure-mode miners (verbosity, keyword-parroting, mid-sentence truncation, generic semantic disagreement) yield 2,336 high-quality preference pairs after cross-encoder filtering. 3,600-step DPO with 95% SFT mixing keeps the gentle anchor while sharpening answer quality
- Cross-encoder DPO judge — the v9r5 cross-encoder now filters preference pairs before DPO training, rejecting cases where the "rejected" answer is just an alternative valid phrasing. Stops DPO from teaching the model to suppress correct-but-different answers
- Mid-training benchmark eval —
train_dpo.py now runs SQuAD + internal benchmarks every 200 steps and saves best.pt by combined EM instead of val loss. Solves the long-standing "DPO val_loss is the wrong signal" issue and lets DPO runs converge without manual snapshot sweeping
- Disambig-pair signal added — the disagreement miner caught residual cases where the model picked a wrong fact (date, name, attribute), adding 677 new pair signals beyond v11's name-echo focus
- v12 retriever live — bi-encoder retrained on shuffled v12 data (fixed an unshuffled-source bug from v11 that left the encoder seeing only the first 10% of the file). New 57.2M-passage FAISS index includes 1.26M synthetic infobox passages and 263K Disambiguation-labeled passages. New cross-encoder trained at bs=48 (vs v11's bs=32) on 6.5M mixed 85/15 easy/hard pairs.
- v11 failure cases fixed — "Who is Marie Curie?" now routes to the actual scientist (not the UK charity); "What is Mercury?" returns the multi-meaning disambig answer; "Who directed Inception?" pulls the right infobox passage and answers "Christopher Nolan."
SFT benchmark (gold-context, 500q): SQuAD EM 17.6% (v11: 17.4%) · Internal v11 EM 60.8% (+16.0pp over v12 SFT baseline) · Internal v12 EM 64.4% (+14.6pp over baseline) · Retrieval passage_match@3 76.0% (v11 stack: 47.6%) · infobox/disambig retrieval went from 0% to 85%+ since v11 didn't have those passage types in its index at all
v11 May 2026
- Disambiguation handling — ambiguous names like "Mercury" or "Phoenix" now return a multi-meaning answer ("Mercury can refer to several things, including the planet Mercury, the chemical element Mercury, or the Roman deity Mercury") instead of being dropped or answered single-sense. Powered by 318K-entry disambig lookup built from Wikipedia disambig pages
- Sense-specific variant questions — sense articles like Mercury (planet) now generate "What is the planet Mercury?" extractive Q&A, giving the model 245K additional disambiguated training examples
- Surname collision acknowledged — "When was Salgado born?" now responds "Salgado refers to multiple distinct entities, including…" instead of guessing one person's birth date
- Fresh Wikipedia corpus — re-extracted with template expansion preserved. Numbers and units in passages now intact: "The Eiffel Tower is 330 m tall" instead of "is tall". Disambig pages kept (332K of them) so the retriever can serve them too
- 2.7× larger SFT dataset — 37M training entries (vs 14M for v10) with sqrt-rebalanced distribution and 1.05M direct disambig Q&A pairs added
- Cleaner Who-is form — DPO trained on 4,660 mined echo failures (3× v10's pool) drives name-echo rate to ~0% on benchmark
- v11 retriever now live — bi-encoder retrained on v11 data (recall@10 96%, +3pp over v9r5 phase 1), new cross-encoder (91.9% hard-discrimination val acc, ~3pp over v10). FAISS index rebuilt at pq_m=128 with fp16 fallback for exact re-rank on top-200 PQ candidates. Index covers 55.7M passages (incl. 263K synthetic disambig passages so "What is Mercury?" can return the multi-meaning answer).
v11 capability benchmark: 96% overall (v10: 40%) · SFT eval/loss: 0.0973 (v10: 0.1107, −12%) · Bi-encoder recall@10: 96% · Cross-encoder hard-val acc: 91.9%
v10 April 2026
- DPO preference tuning — Direct Preference Optimization teaches the model to prefer role descriptions over echoing entity names. Internal EM jumped +20pp over v9r5
- Two-stage retrieval (v10) — bi-encoder (bs=88, recall@10 97.7%) finds top-20 candidates, cross-encoder (100% golden QA accuracy) re-ranks by scoring question+passage together. Both retrained on v10 data with new question types
- Longer definitional answers — fixed over-trimming that reduced "An American politician who served as the 44th president" to just "An American politician"
- Causal Why framing — Why answers now start with "Because" / "Due to" instead of bare noun phrases
- New question types — physical properties (boiling point, half-life), dimensions (area, elevation, "How tall is X?"), composition ("What is X composed of?"), and list-style answers ("What are the Great Lakes?")
- Sqrt question-type rebalancing — no question type gets starved of training signal
Internal EM: 58.4% · Internal F1: 85.0% · Golden QA EM: 43.4% (53 questions) · SQuAD EM: 16.6%
v9r3 April 2026
- Section & paragraph metadata — passages tagged with section name, section number, and paragraph position
- Natural-language metadata format — switched from § structural tokens to plain English
- SQuAD-style templates — "Which X" superlatives, definite-article swap questions
- Garbled-question detection — new audit check catches regex-artifact fragments
Val loss: 0.1048 · Internal EM: 40.8% · Internal F1: 66.8%
v8 April 2026
- Disambiguation articles blocked — "Steve Jobs (book)" no longer generates "Who is Steve Jobs?"
- Nickname matching expanded — 150 nickname→formal name mappings (Jeff/Jeffrey, Tim/Timothy, etc.) so definitional questions generate for all common name variants
- Definition cleaning fixed — Batman, Superman, Spider-Man no longer reduced to just "superhero"
- Domain identifiers preserved — C3.ai, Tesco.com no longer split by the text normalizer
- Data quality audit script — automated 7-check quality gate catches regressions; title mismatches dropped 179→3
- Ambiguous questions (Mercury, Victoria, etc.) dropped instead of generating confusing training signal
- v8 retriever: phase 1 encoder already outperforms v5r5 deployed by 2.8x on gold rank
SQuAD EM: 19.0% (new high) · Internal EM: 47.4% (new high) · Internal F1: 73.7% (new high, +3.3 over v6r2)
v7 April 2026
- Definitional Q&A restricted to own article — "Who is Barack Obama?" only from the Barack Obama article, not tangential articles
- Generic subject filtering — "The amended results", "The following squads", "Her next project" no longer used as question subjects
- Pronoun resolution improved: "He in turn employed" → resolves to article subject
- Catalog numbers no longer extracted as years (e.g., "PKS 1614+051")
- Historical verb tense in "What does X do?" answers
- Quality audit: 78% good examples (up from 70%)
v6r2 April 2026
- Title in training context — model now sees "Context: Title — passage text" matching retriever format
- Fixed "Context:" prefix mismatch between training and inference
- Fixed prompt order (user→system→assistant) in inference scripts
- Optional first-paragraph boosting for retrieval re-ranking
- 15K training steps (down from 25K — model plateaus at ~11K)
SQuAD EM: 18.8% (+92% over v6) · Internal F1: 70.4% (best ever)
v6 April 2026
- Expanded question extraction: "What did X [verb]?", "What type of?", "Why?" patterns
- Rebalanced span-answer ratio (fewer terse extractions)
- SFT context truncation fix (keep beginning of paragraph, not end)
- Behavioral examples added (greetings, identity, refusal) — too few to have effect (51 in 25M)
- Data quality improvements: better subject rejection, definition filtering
SQuAD EM: 9.8% · Internal EM: 46.4% · Internal F1: 68.2%
v5 March 2026
- 27M training examples with augmented rephrasings and typos
- "How many" counting questions, action questions ("What did X do?")
- Deployed to Modal (serverless T4 GPU) at gpt.shult.ca
- Dense retrieval with MLM-pretrained bi-encoder (124M params, 768d/12L)
- Paragraph-aligned passage index with title prefixes (41.7M passages)
SQuAD EM: 10.8% · Internal EM: 45.6% (gold context)
v4 March 2026
- Diverse question types: When, Why, What does X do, awards, birth dates, acquisitions, population
- Span-extraction post-processing (short answers alongside full sentences)
- Article title context for Q&A generation
- 18.9M training examples (up from 4.5M)
SQuAD EM: 7.8% (+39x over v3)
v3 March 2026
- Major data quality overhaul: stricter grounding, anti-parrot filtering
- Context-aware SFT training (system message with passage)
- 520M base model (PPL 14.6)
- Dense retrieval Phase 1+2 (384d/6L bi-encoder, 30M params)
Data quality was the biggest win: +42.8% EM over v2
v2 September 2025
- Improved extraction: better subject resolution, answer cleaning
- Still limited to "What is X?" / "Who is X?" questions
- 124M model
v1 August 2025
- Initial Q&A pipeline
- "What is X?" / "Who is X?" patterns only
- 124M model, q-only training (no context)
- TF-IDF retrieval baseline