# Earning The Right To Fine Tune

*What a fund-accounting hack (re)taught me about the smallest available fix —  and where that thread leads next.*

By [papajams.eth](https://paragraph.com/@papajams.eth) · 2026-09-06

ai, fine-tuning, llm, models, qwen, thinkingmachines, ylookup, encode

---

At the ILOAI 2026 machine-translation olympiad, I fine-tuned a 7B model on chain-of-thought examples and got 0.075, below the untouched 14B model's 0.123. The eventual breakthrough wasn't another training run. It was a parser bug: a heuristic meant to strip leaked reasoning was also deleting correct answers that happened to start with the same words. One bug fix mattered more than every fine-tuning attempt combined. I finished second out of 48, a few hundredths behind first, on a T4, solo.

![](https://storage.googleapis.com/papyrus_images/253d51eadb03080d36515a2cb99e3a5b1be7633264f692d8fc53f62d87203277.png)

[ratiocine.trustfall.xyz](http://ratiocine.trustfall.xyz) — built a linguistics game out of the experience

This weekend I entered a chess-engine hackathon — build a bot from scratch, CPU-only, no borrowed engines. Every search improvement I added — transposition tables, null-move pruning, principal variation search — made the bot worse.

The cause wasn't the search. It was two one-line bugs: the clock was being read as seconds when the platform sent milliseconds, so the bot moved in 0.1 seconds instead of 4; and the value network's output was signed for White regardless of whose turn it was, so playing Black meant optimizing for the opponent. Fixing both took the bot from zero wins in four games to a real, climbing rating — before a single architecture change. Same lesson, sharper edge: the bugs didn't just cap the upside, they turned every subsequent improvement into a step backward.

![](https://storage.googleapis.com/papyrus_images/788cb193e4b64bdaf7d373064f9d243f1572b4726e6c0e3af441cf660fd75d62.png)

[https://farcaster.xyz/papa/0x036f5437](https://farcaster.xyz/papa/0x036f5437)

> The ceiling is often somewhere around the model, not inside it. The harness, parser, decode config & deterministic boundaries can matter more than another round of training— and, as this weekend made clear, so can how rigorously you audit that harness itself.

And I don't think that's a hackathon-specific observation.

A shift in AI right now is from models that can do things to systems that can reliably deploy those capabilities. The model is becoming a component rather than the product.

![](https://storage.googleapis.com/papyrus_images/983a3d8487607542a901e22fe5ef359bb48c947b57a10045fd851fdc2b7edfa3.jpg)

### Where this points next

SpreadsheetBench is 400 tasks graded against a golden workbook: a neat, bounded version of a problem that's much messier in the wild.

A small business has actual books, actual bank records and actual invoices, often without a fund's back office to reconcile them. Private-markets and back-office work like this is a multi-trillion-dollar wedge — reconciliation and reporting that funds and small businesses alike still do largely by hand. That's what I'm building with Sikizana: an AI bookkeeper that sits on top of Xero, chases overdue invoices, matches receipts to transactions, and cites the actual HMRC rule before it touches a ledger.

![](https://storage.googleapis.com/papyrus_images/17173ee993c5ed806eaa9d0148976c188e96ac6188d08c3d59539a5cdd31dc6e.png)

[https://sikizana.persidian.com](https://sikizana.persidian.com)

Xero and a growing crop of AI bookkeeping tools can already extract documents, categorize transactions & suggest reconciliations. The interesting problem is no longer whether an LLM can read a receipt. It's what happens when the answer is wrong — & whether the system can show you why it made the decision, what rule supported it, and who approved it.

That's the same principle tieout exposed at the level of a single graded cell, and the same one the 68-to-88 gap exposed at the level of the whole harness: find the discrepancy, show your work, test it against itself, and don't let the model post the journal entry — or the fine-tune claim — without a human able to verify how it got there. The infrastructure has to make all four rungs of that ladder the default, not the thing a deadline lets you skip: fast enough evals to trust a result, output structured enough to fail loudly, settings tested against each other, and reasoning spent deliberately rather than maxed out because nobody checked the cost.

If task-specific training is where the frontier is heading, then the opportunity underneath it isn't just training better small models. It's building the environment those models have to operate inside — one that makes the last twenty points as available as the first were, without needing a leaderboard to force the discipline.

The question becomes less _can the model do this?_ and more _what is it allowed to do, what evidence does it have to provide, and can you prove it did it right?_

That is the part of AI infrastructure that interests me most now: not making the model smarter for its own sake, but turning intelligence into something a business can actually trust.

The cell was always going to be right or wrong.

What mattered — what still matters, twenty points later — was whether you could tie it back to where it came from.

* * *

Repo: [github.com/udirobert/tieout](https://github.com/udirobert/tieout) · Related: [Get the Fridge Fixed](https://medium.com/@ungethe/get-the-fridge-fixed-b3056bea6674), [What a GPU Doesn’t Understand About Language](https://medium.com/@ungethe/what-a-gpu-doesnt-understand-about-language-bccacf3654c6), [Building Small](https://medium.com/@ungethe/building-small-9aea8bf5236e)

---

*Originally published on [papajams.eth](https://paragraph.com/@papajams.eth/earning-the-right-to-fine-tune)*
