I entered an AI olympiad for endangered languages. Here's what nearly broke it — and what won it second place.

Every year, the International Linguistics Olympiad hands human competitors a page of examples from a language they've never seen — Ubykh, Coastal Marind, Plains Cree, some with fewer than a hundred living speakers — and a few hours to reverse-engineer the pattern underneath. This year, for the first time, they let AI models enter too. Submit a script and a set of weights, and the organizers drop it onto a sandboxed T4 GPU: 16GB of memory, thirty minutes on the clock, no internet, no second chances mid-run.
I entered solo, last minute. You theoretically got ten submissions per day through the competition, I managed eight, and most of what I learned came from watching things fail in instructive ways.

The first day went to the obvious move: fine-tune a model on practice problems. I had a frontier model generate chain-of-thought reasoning for a few hundred examples, trained a 7B model on top of it, and expected an easy win. It scored 0.075. The untouched 14B base model, no training at all, scored 0.123. A few hundred worked examples simply couldn't out-argue everything a bigger model had already absorbed in pretraining — fine-tuning is a bet that your data outweighs the model's existing knowledge, and against a 7B-to-14B gap, that bet doesn't clear. I dropped the fine-tuned model entirely.

The single highest-return fix of the whole project, though, wasn't a model change at all. It was a bug in my own answer parser, sitting there for ages: any line starting with "The," "We," or "There" got filtered out, on the assumption it was leaked reasoning rather than an answer. Except sometimes the correct translation is "The sun rises." The filter had been quietly deleting correct answers the entire time, and nothing about the score would have told me that directly — it just looked like the model wasn't very good. Removing that one heuristic did more than any prompt rewrite that followed.
Prompting itself was mostly an argument with myself that I kept losing. Generic instructions plateaued around 0.12 no matter what I tuned around the edges; writing five separate prompts, one per task type, pushed the same model to 0.235 — exact match doesn't reward being close, it rewards being exactly, structurally right. I tried telling the model to answer "verbatim only, no extra words," expecting cleaner output, and watched accuracy drop by nearly half across three submissions because the model started truncating answers it shouldn't have. Three submissions gone, out of a budget of ten, over one sentence of instruction.
What actually worked was closer to the opposite instinct: let the model reason out loud, but only where it needed to. Wrapping reasoning and answer in separate tags, applied only to the two task types that genuinely required judgment — translation and fill-in-the-blank — while letting the more mechanical tasks answer directly, got me to 0.1255 locally. Full verbose reasoning on every task would have taken forty-five minutes; the trick was knowing which four of the five task types didn't need it. On top of that sat a three-tier time guard: if the run fell behind schedule, it shrank the reasoning budget; if it fell badly behind, it dropped straight to short, direct answers for whatever was left. A model that reasons beautifully and then runs out of clock is worth exactly as much as one that never reasoned at all.

The final standings went up while I was still writing about this. Forty-eight teams, scored on a hidden private split none of us had seen. I came second — 0.192, a hundredth of a point behind first — solo, on a single T4, with a parser bug and three burned submissions along the way.

The full write-up has the actual scores at each stage, the tooling that made the iteration loop fast enough to try all this in the first place, and a few hard-won rules (float16 not bfloat16 on a T4; absolute paths only; don't let a newer library touch your tokenizer before you push it) that would've saved me a day each if I'd known them going in.


