The domain imports the application layer. The engine identifies the forbidden edge in order.ts:1.
Inspect the source change
The packaged comparison changes the files below. It is not a one-line-only causal experiment.
packages/app/checkout.ts
Conforming fixture
import { priceOrder, type Order } from '../domain/order.js';
export function checkout(order: Order): number {
return priceOrder(order);
}
Drifted fixture
import { priceOrder, type Order } from '../domain/order.js';
export function checkoutLabel(total: number): string {
return `checkout:${total}`;
}
export function checkout(order: Order): number {
return priceOrder(order);
}
48 typed questions. Open weights on a laptop against a hosted API.
Accuracy / Laya, localRED · 39 / 48
Accuracy / Jev, hosted46 / 48
Median time / local vs hosted, incl. network9 ms vs 378 ms
Same answer as the reference code48 / 48
The local model needs no key and matches its reference on every row. It is 14.6 points less accurate than Jev here, beyond the 10-point limit set before the run.
Measured 29 Sep 2026, judged against the record as published
The claim is refuted, by the criterion jev-vs-best-baseline. A linter written from the same rules (heuristic-lint) missed 0 of 30 drift items; Jev missed 1 of 30. On rules a linter can express, a fast model gate added nothing that linter did not already catch. By one item: a difference of 3.3%, 95% interval -8.3%–16.7%. Jev's only miss, c026, is a comment-only RED item, which the pre-registration itself describes: "Fair under the rules, but not architectural drift in the usual sense."
Jev · missed drift1 of 30 (3.3%)
Linter from the rules · heuristic-lint0 of 30 (0.0%)
Jev · false reject7 of 30 (23.3%)
Cascade · sent to the reviewer26 of 60 (43.3%)
Mean cost per change · mixed-basis: Jev listed price + reviewer API-equivalent$0.2743 cascade · $0.6495 reviewer alone
LLM reviewer · spotlight bar, restricted setupbar met, spotlight held · 0 of 90 missed · diff-blind in 131 of 180 runs
The five criteria, as pre-registered
Jev misses more than 10% of drift (RED items it accepts or abstains on). passes, not established at this N · 1 of 30 (3.3%), 95% interval 0.6%–16.7%
Jev falsely rejects more than 25% of clean changes (GREEN items it rejects or abstains on). passes, not established at this N · 7 of 30 (23.3%), 95% interval 11.8%–40.9%
The cascade misses more drift than the reviewer alone, on the same RED items (an abstention counts as a miss for both). passes, not established at this N · 0 of 30 (0.0%) against 0 of 30 (0.0%); difference 0.0%, 95% interval -11.4%–11.4%
Jev misses more drift than the better model-free baseline, on the same RED items (an abstention by Jev counts as a miss). refuted · 1 of 30 (3.3%) against 0 of 30 (0.0%); difference 3.3%, 95% interval -8.3%–16.7%
The cascade's mean cost per change is not below 50% of the reviewer-alone mean cost per change, on the same items and the pinned price bases (Jev at its listed price, the reviewer at the API-equivalent total_cost_usd it records). passes · $0.2743 against $0.6495 per change, ratio 0.422 (mixed-basis: Jev listed price + reviewer API-equivalent)
Pre-registered · amended 28 Sep 2026, before any counted gate run. The LLM reviewer is re-pinned from nina 0.28.10 to 0.34.0, and a bar nina's reviewer must clear for a spotlight is fixed in advance. Amendment 01, the record · What changed
Amended again 28 Sep 2026, before any counted run. The reviewer's file and git tools are fenced to its workspace after the dry run's isolation probe read a file outside it. Amendment 02, the record · What changed
175 rules that the plugins ship (145 primary, 30 sampled) and 13 controls (6 written for this experiment, 7 from EXP 005), each translated blind into a static checker's constraints
Median expressible share0.0% (bar 25.0%)
Rules expressible, primary stratum1 of 145
Partial (a model still needed)86
Controlscalibration bar met
The median expressible share across the 7 plugins with rules is 0.0%, below the pre-registered 25.0% bar: the premise is refuted: stage 2 is cancelled and the census is the result.
Pinned census gate · results.json sha256 26eb7287e88f · The results record
Station evidence is unavailable. Follow the station’s full experiment link, or inspect the source.
Four ways to question a result, one decision model you can run yourself, and one experiment written down before it ran. These machines are a conceptual map; the records are real runs of authored fixtures, one measured model comparison and one pre-registered experiment, measured 29 Sep 2026: the claim is refuted. Selecting a station reads evidence. It does not run an agent or a model.
Try to prove it wrong.
A useful check has to reject a plausible failure. Inspect the input, the result and the limits.
Every experiment on this site, and where it stands.
It lists each experiment whose pre-registration this site publishes; a record published only in the repository joins the list once its site copy does. EXP 001–003 show their recorded run: its date and the two scores, as published in the run's record. Every later experiment shows the status sentence from its own field note, unedited, or says it has no note yet. A refuted claim stays refuted.
Packaged fixtures, run with BCE 0.3.1. This tests the mechanism; it does not measure agent productivity or production reliability. Recording environment: GitHub Actions. Selectors inspect the recording; reproduction runs on your machine.
Then a decision you can own.
A second bench, because this one compares two models rather than one rule on two trees: open weights on a laptop against a hosted API, on the same questions.
BENCH 02 / DECISIONS
EXP / 004Laya vs Jev: typed decisions without lock-in
Can an open model answer the same typed decisions on a laptop, with no vendor key?
WHAT IS A SYSTEM-1 DECISION MODEL?
It answers a typed question about a situation: pick one of these options, give a score, or say yes or no. It returns that answer with a confidence. It does not write text, so software can branch on the answer directly.
WHY NO VENDOR LOCK-IN MATTERS
A decision inside your control flow is a dependency. Jev is a hosted API: every answer needs its key, its network and its price. Laya publishes open Apache-2.0 weights that run on your own machine, so the hosted model can become a comparison instead of a requirement.
SPECIMEN A / LAYA · OPEN WEIGHTS, LOCAL39 / 48correct · median 9 ms per decision
SPECIMEN B / JEV · HOSTED API46 / 48correct · median 378 ms, including the network
Result: the local port gives the same decision as the model’s reference code on 48 / 48 rows, and its median call here took 9 ms on the laptop against 378 ms for the hosted call including its network round trip, but it is 14.6 points less accurate than Jev. That is beyond the 10-point limit set before the run, so the accuracy check is RED and stays that way.
REDWithin 10 points of Jev. 39 / 48 against 46 / 48: 14.6 points behind.
GREENConfident answers are right. 2 of 2 answers at confidence ≥ 0.9 were correct, but only 2 of 48 reached it.
GREENSame answers as the reference code. 48 / 48 same decision; largest probability difference 0.0021 (limit 0.01).
GREENFaster than the hosted call, network included. Median 9 ms on the laptop against 378 ms for the hosted call, including the network.
REPRODUCE / APPLE SILICON + UV · FROM A CLONE OF THIS REPOSITORY
48 hand-written rows on one busy laptop (load average 36–47 during the run) is a mechanism check, not a benchmark. It does not predict accuracy on your decisions. Jev’s latency includes the public internet; Laya’s confident-answer figure rests on 2 answers. Opening this page runs nothing; the browser demo downloads and runs Laya only when you ask it to.
The check is only the beginning.
These experiments make failure visible in small, authored examples. They do not establish customer savings, independent replication or how an agent will perform on unseen work.
Our internal build note follows the gap between a green pipeline and a verified user journey. An author-reported account, with its private evidence boundary stated.
An open-weight typed-decision model that runs on your own device: give it a situation and a choice, score or yes/no question, and read the probability behind each answer. No server, no key. We measured it on a laptop against a hosted model and published where it fell short.
Laya is by ConvAI Innovations; the browser conversion is by an independent author. Odin checked provenance and agreement with the original on three fixed cases.
nina composes the agents, hooks and checks that run Claude Code on a project. We measured one part of it, its reviewer at 0.34.0, deciding whether a change breaks mechanical architecture rules, with a bar fixed before any counted run: 1 of 90 runs missed drift (passes, not established at this N on the items), 2 of 90 runs falsely rejected (passes on the items), the same verdict on 57 of 60 changes, no local patch to nina, 0 of 180 harness failures, and in 177 of 180 runs a tool output showed it at least one line of the change. Nothing else about nina was measured.
nina is by Marcos Schulz (xhulz), github.com/xhulz/nina — used with the permission of its author, as confirmed by Odin Labs. The upstream fixes above were opened by Odin Labs.
Seeded defects, fixture trees, recorded gate runs and a standalone chain verifier. Inspect the research materials behind the blueprint-conformance paper.
Private collaboration For internal teams, customers & early adopter partners.
Factory Intelligence
Understand what good work costs—and where to improve next. Shape Factory Intelligence around the quality, cost and outcomes that matter to your team. See the horizon and where we stand today.
Give agents a place to practice on work that matters. Shape Odin Gym around your workflows, with isolated tasks and verifiable outcomes as the design goal.
Developed for Odin’s internal use, private customers and early adopter partners. Bring your use case to discuss scope and early access. Become an early adopter partner ↗