Project 05 · Retrieval · 2026In progress · Phase 3
Does grounding still beat a frontier model in Korean?
A benchmark comparing grounded retrieval against frontier models answering from parameters alone, over Korean corporate filings from DART. The English-language version of this question is close to settled. The Korean one is not, and the gap is where the interesting answer lives.
[TODO: one paragraph on method — multilingual-E5 for embeddings, dense-only retrieval, and the Mr. TYDI Korean Wikipedia notebooks that justify both choices. Say plainly that those two decisions were made and frozen before the DART results existed.]
[TODO: what you expect to find, written now, before you have the numbers. A pre-registered guess you can be wrong about in public is worth more than a clean result explained afterwards.]
RoleSole author
CorpusDART corporate filings
BaselinesGPT-4o, Claude, Gemini
StatusTest set in construction