Cover photo

A slime mould is not an agent

Five experiments tested whether a slime mould's network rule helps AI models decide, route and retrieve, and it never helped overall.

Physarum polycephalum is a single-celled slime mould with no brain, and it still builds good networks. Put oat flakes on a map of the Tokyo region and it grows tubes between them that resemble the rail system (Tero et al., Science, 2010). The earlier model by Tero, Kobayashi and Nakagaki (2007) writes the behaviour down as one rule: flow runs through tubes like current through wires, tubes that carry a lot of flow thicken, and unused tubes shrink. On a network with one start and one goal, that rule provably converges to the shortest path. Astronomers later used a Physarum-inspired model to trace the filaments of the cosmic web between galaxies (Burchett et al., 2020).

So we asked a practical question. If a slime mould explores a problem before an AI model acts on it, does the model do better?

We built the rule as a Rust engine first. It gives byte-identical results natively and in WebAssembly, and it finds the textbook shortest path on 256 of 256 random test networks. That engine ran the first two experiments. The third used a separate implementation of the same rule, which we later re-ran on our engine and matched to within about two parts in a hundred billion. The fifth used a separate implementation of the same flow rule, and the fourth a Physarum-inspired, reward-driven variant built for routing. The rules and the pass mark were fixed before each deciding run, and three of the five included a control that did the same flow calculation without the adaptive growth.

Mold evidence in the prompt.

On 40 hard public questions from JevBench, we added the mold's network summary to the prompt of Jev, a decision model. It did not help. This is an independent experiment on the public split, not a ranked JevBench score, and the model's run-to-run variation made the official verdict inconclusive.

A network we built ourselves.

Next we removed the extraction step. We wrote 120 synthetic career-planning personas, graded by career rules of our own, and built each one's network directly from its typed profile. We then compared the model given a one-step flow summary with the model given the same summary plus the mold's adaptive result. The adaptive result made decisions clearly worse, by 0.16 nats of log-loss, in all six repeated passes. The flow summary itself helped over the plain network, but mostly because it showed options already ruled out by hard constraints as zero, and those checks overlapped with our own grading rules. That part is partly by construction.

An independent answer scorer.

On all 231 public JevBench questions, the mold scored the answer options on its own, and we fused its scores with Jev's answers. The mold carried real signal, about 20 points above chance, but it added nothing Jev did not already know: every fusion mode failed, and that fusion step was the part tested blind. The mold-only result and the adaptive-versus-static comparison had already come out of the spec's own development run on the same questions, and our run reproduced them: the adaptive version scored the same as a single static flow calculation. Again, an independent experiment on the public split, not a ranked score.

A router between language models.

In a pilot (five seeds, one budget) on the public ParetoBandit routing benchmark, the mold routed requests between three models while quality, availability and capacity changed underneath. It lost to a standard discounted bandit under an outage and under a capacity cap, and it won by a hair under a quality drop. It did recover more fully once a shock ended (94 against 85 percent of its earlier reward), but it took more steps to adapt during the outage.

A builder of structure.

Finally, the role where Physarum has real successes: growing a network over about 27,000 text passages from HippoRAG's multi-hop benchmarks, so that evidence spread across several passages becomes easier to find. The mold network kept only 70 to 91 percent of the direct links between evidence passages that a plain nearest-neighbour graph keeps. Its best gain was about one point of recall at five, against a pass mark of seven. Against the same flow calculation without adaptation, the difference was zero on two datasets and one point, not significant, on the third.

Where it works, and where it does not

The engine works. It finds shortest paths, it is deterministic across platforms, and it is fast when one route clearly wins. What did not work is the thing we were testing: in the three experiments that had a static-flow control, the adaptive growth never beat a single flow calculation, and once it was clearly worse. In the other two, the mold did not beat ordinary baselines overall. In those three controlled tests, one flow calculation already held everything the adaptation found.

The limits are real. The persona and routing tests were small, the routing test was a pilot, and our answer keys in the persona test were rule-based rather than real outcomes. We also tested the flow-network family of Physarum models. The particle-based model behind the cosmic-web work is a different model, and these results do not rule it out. The upgrade path, if anyone wants one, is to test that model against the same static controls.

We stopped the project. We kept the engine and the five reports, and we would rather publish a clean negative result than build a product on a beautiful idea.

A product that did ship

Not every idea ends in a negative result. Tideproof is a free scan that tells you in seconds how exposed your job is to AI. Describe your work in a few sentences, or drop in your LinkedIn PDF export, which is parsed in your browser and never uploaded. You get a 0 to 100 exposure score, a band, and three building blocks to start with, each with one concrete first move. It has nothing to do with the slime mould: the scoring is plain code, and its weights are published on the methodology page. It is education, not financial advice. Try it, and tell us where it gets you wrong.

Contact the author

I write about AI systems, decision engines and honest measurement as metaend. If you want to reach me, build on any of this, or tell me where it breaks, the door is here:

metaend on Quilibrium

More writing lives at paragraph.com/@metaend.

Written by metaend.