What should an agent’s
work have to prove?

We build the checks around the work.
Open a machine to inspect an authored experiment and its recorded result.

Discuss a workflow

Open a station

Odin software factory, conceptual assembly drawingAn isometric factory connects a blueprint drafting table, robotic workcell, conformance gate, evidence records, a decision console, a triage bench and a floor inspection bench. Use the seven station links below to explore each part.INTENT → EVIDENCEFLOOR 00001060203040507ODIN / RESEARCH & DEVELOPMENTORTHOGRAPHIC ASSEMBLY / CONCEPT DRAWINGDWG. OD-000 / ABLUEPRINT → WORKCELL → GATE → RECORDNOT TO SCALE
Concept drawing. Separate experiments.Odin / R&D
01 / BlueprintRecorded experiment

Can a dependency cross the line?

An unchanged architectural rule

domain-cannot-import-app
domain → app: forbidden
Conforming treeGREEN / 100
Reverse importRED / 60

The domain imports the application layer. The engine identifies the forbidden edge in order.ts:1.

Inspect the source change

The packaged comparison changes the files below. It is not a one-line-only causal experiment.

packages/app/checkout.ts

Conforming fixture

import { priceOrder, type Order } from '../domain/order.js';

export function checkout(order: Order): number {
  return priceOrder(order);
}

Drifted fixture

import { priceOrder, type Order } from '../domain/order.js';

export function checkoutLabel(total: number): string {
  return `checkout:${total}`;
}

export function checkout(order: Order): number {
  return priceOrder(order);
}

Conforming source / Drifted source

packages/domain/order.ts

Conforming fixture

export interface Order {
  total: number;
}

export function priceOrder(order: Order): number {
  return order.total;
}

Drifted fixture

import { checkoutLabel } from '../app/checkout.js';

export interface Order {
  total: number;
}

export function priceOrder(order: Order): number {
  return checkoutLabel(order.total).length + order.total;
}

Conforming source / Drifted source

Read the unchanged blueprint
{
  "apiVersion": "blueprint-conformance/v1alpha1",
  "kind": "EngineeringBlueprint",
  "metadata": {
    "id": "typescript-module-layering",
    "name": "TypeScript module layering",
    "version": "0.1.0",
    "status": "approved",
    "ownerRole": "architecture-owner",
    "stewardRole": "blueprint-steward"
  },
  "intentRefs": [
    "docs/typescript-module-graph.md#directional-layering"
  ],
  "scope": {
    "repositories": [
      "example-org/storefront"
    ],
    "paths": [
      "packages/**/*.ts"
    ]
  },
  "architecture": {
    "components": [
      {
        "id": "typescriptModule",
        "type": "typescriptModule"
      }
    ],
    "relationships": [
      {
        "from": "packages/domain/**",
        "to": "module:packages/app/**",
        "type": "imports",
        "allowed": false
      }
    ]
  },
  "constraints": [
    {
      "id": "module-surface-exists",
      "type": "requiredComponent",
      "severity": "critical",
      "component": "typescriptModule"
    },
    {
      "id": "domain-cannot-import-app",
      "type": "forbiddenDependency",
      "severity": "critical",
      "from": "*",
      "to": "module:packages/app/**",
      "scopePaths": [
        "packages/domain/**"
      ]
    }
  ],
  "evidenceRequirements": [
    {
      "type": "staticAst",
      "required": true,
      "onMissing": "block"
    }
  ],
  "approvals": [
    {
      "role": "blueprint-steward",
      "stage": "ratify"
    }
  ],
  "minEngineVersion": "0.3.0",
  "extraction": {
    "profile": "typescript-module-graph",
    "paths": [
      "packages/**/*.ts"
    ],
    "minFiles": 2
  }
}
Exact released blueprint

BCE 0.3.1 · 2026-10-11
Recording, source hashes and environment

Open full experiment

Four ways to question a result, one decision model you can run yourself, and one experiment written down before it ran. These machines are a conceptual map; the records are real runs of authored fixtures, one measured model comparison and one pre-registered experiment, measured 29 Sep 2026: the claim is refuted. Selecting a station reads evidence. It does not run an agent or a model.

Try to prove it wrong.

A useful check has to reject a plausible failure. Inspect the input, the result and the limits.

Every experiment on this site, and where it stands.

It lists each experiment whose pre-registration this site publishes; a record published only in the repository joins the list once its site copy does. EXP 001–003 show their recorded run: its date and the two scores, as published in the run's record. Every later experiment shows the status sentence from its own field note, unedited, or says it has no note yet. A refuted claim stays refuted.

  1. EXP 001Can a dependency cross the line?RecordedRecorded 11 Oct 2026: the same rule scored the conforming tree GREEN / 100 and the drifted tree RED / 60.
  2. EXP 002Does the boundary hold in Python?RecordedRecorded 11 Oct 2026: the same rule scored the conforming tree GREEN / 100 and the drifted tree RED / 60.
  3. EXP 003What happens when configuration widens?RecordedRecorded 11 Oct 2026: the same rule scored the conforming tree GREEN / 100 and the drifted tree RED / 60.
  4. EXP 004System-1 decisions without lock-in: Jev vs open-weight LayaMeasuredMeasured on one laptop, 48 authored rows.
  5. EXP 005Jev as a fast gateResults published · 2 amendmentsPre-registered; measured 29 Sep 2026: the claim is refuted.
  6. EXP 006nina reviews the changeResults published · 1 amendmentPre-registered; measured 01 OCT 2026: bar met, nina in the spotlight. Amended 01 Oct 2026, before any counted run.
  7. EXP 007Which of your gate's rules need a model at all?Results published · 1 amendmentPre-registered; measured 7 Oct 2026: the premise is refuted.
  8. EXP 008Hand over the cache, not the textPre-registered · 2 amendmentsPre-registered; not yet run.
  9. EXP 009Unload leaves no traceResults publishedP1 refuted: 1 of 400 H1 trials diverged from R0 (95% Wilson upper bound 1.40%). P2 supported: 0 silent-inert and 0 false-inactive kernel configurations of 120. All 5 validity gates passed, so the record is informative.
  10. EXP 011Where does a model reviewer earn its keep?Pre-registeredPre-registered; not yet run.

Start with an architectural boundary.

Three published BCE recipes. The same rule, applied to a conforming tree and a deliberately drifted tree.

SELECT EXPERIMENT
Watch the public runs
EXP / 001RECORDED EXPERIMENT

Can a dependency cross the line?

A domain module imports the application layer. Does the same blueprint distinguish it from a conforming tree?

SPECIMEN A / CONFORMINGGREEN / 100The authored boundary holds.
SPECIMEN B / DRIFTEDRED / 60A deliberate violation is detected.

One unchanged rule. Two different trees. The experiment succeeds when the engine tells them apart.

Read the engine output
recipe module-layering [TypeScript/JavaScript · direct module graph] — Keep domain code below application code
GREEN conformant: score 100, exit 0
RED drift: score 60, would exit 1, violation domain-cannot-import-app
  observed forbidden direct import module:packages/domain/order.ts -> module:packages/app/checkout.ts is present
  evidence packages/domain/order.ts#L1
bce demo: module-layering discriminates GREEN from RED
REPRODUCE / NODE 22+
npm exec --yes --package=bce-engine@0.3.1 -- bce demo --recipe module-layering
RECORDED 2026-10-11T05:47:52.210ZFull result & provenance

Packaged fixtures, run with BCE 0.3.1. This tests the mechanism; it does not measure agent productivity or production reliability. Recording environment: GitHub Actions. Selectors inspect the recording; reproduction runs on your machine.

Then a decision you can own.

A second bench, because this one compares two models rather than one rule on two trees: open weights on a laptop against a hosted API, on the same questions.

BENCH 02 / DECISIONS

EXP / 004Laya vs Jev: typed decisions without lock-in

Try it in your browser
EXP / 004ACCURACY / RED · 3 OF 4 CHECKS GREEN

Can an open model answer the same typed decisions on a laptop, with no vendor key?

WHAT IS A SYSTEM-1 DECISION MODEL?

It answers a typed question about a situation: pick one of these options, give a score, or say yes or no. It returns that answer with a confidence. It does not write text, so software can branch on the answer directly.

WHY NO VENDOR LOCK-IN MATTERS

A decision inside your control flow is a dependency. Jev is a hosted API: every answer needs its key, its network and its price. Laya publishes open Apache-2.0 weights that run on your own machine, so the hosted model can become a comparison instead of a requirement.

SPECIMEN A / LAYA · OPEN WEIGHTS, LOCAL39 / 48correct · median 9 ms per decision
SPECIMEN B / JEV · HOSTED API46 / 48correct · median 378 ms, including the network

Result: the local port gives the same decision as the model’s reference code on 48 / 48 rows, and its median call here took 9 ms on the laptop against 378 ms for the hosted call including its network round trip, but it is 14.6 points less accurate than Jev. That is beyond the 10-point limit set before the run, so the accuracy check is RED and stays that way.

  • REDWithin 10 points of Jev. 39 / 48 against 46 / 48: 14.6 points behind.
  • GREENConfident answers are right. 2 of 2 answers at confidence ≥ 0.9 were correct, but only 2 of 48 reached it.
  • GREENSame answers as the reference code. 48 / 48 same decision; largest probability difference 0.0021 (limit 0.01).
  • GREENFaster than the hosted call, network included. Median 9 ms on the laptop against 378 ms for the hosted call, including the network.
REPRODUCE / APPLE SILICON + UV · FROM A CLONE OF THIS REPOSITORY
cd experiments/laya-vs-jev
uv venv -p 3.12 .venv
VIRTUAL_ENV=$PWD/.venv uv pip install -r requirements.txt
.venv/bin/python run.py            # add --out <file> to avoid overwriting today's recording
MEASURED 2026-09-24 13:34 UTC · APPLE M5 MAX · MLX 0.32.2Full result & provenance Read the field note

48 hand-written rows on one busy laptop (load average 36–47 during the run) is a mechanism check, not a benchmark. It does not predict accuracy on your decisions. Jev’s latency includes the public internet; Laya’s confident-answer figure rests on 2 answers. Opening this page runs nothing; the browser demo downloads and runs Laya only when you ask it to.

The check is only
the beginning.

These experiments make failure visible in small, authored examples. They do not establish customer savings, independent replication or how an agent will perform on unseen work.

Tools leaving
the factory.

Take them apart.
Make something with them.

003ARTIFACTS

The evidence workbench

Seeded defects, fixture trees, recorded gate runs and a standalone chain verifier. Inspect the research materials behind the blueprint-conformance paper.

An earlier research archive. Its engine-availability notes predate the public BCE release linked above.

Type
Research archive
Checked
07 SEP 2026
License
Apache-2.0
Access
Public artifacts
000WORKSHOP

This floor is a project, too.

The drawings, the website and the experiment runner are here to inspect. Propose an experiment, reproduce a result, or help shape the next machine.

Revision
v0.2.0
Type
Public workshop
License
Apache-2.0
Access
Open source

Better outcomes.
Stronger agents.

Private collaboration
For internal teams, customers & early adopter partners.

Odin Gym

Give agents a place to practice on work that matters. Shape Odin Gym around your workflows, with isolated tasks and verifiable outcomes as the design goal.

FOCUS / REAL WORK · PRACTICE · EVALUATIONExplore Odin Gym for your team

Developed for Odin’s internal use, private customers and early adopter partners. Bring your use case to discuss scope and early access. Become an early adopter partner

Notes from the floor.

What happened.
What it means. What comes next.

EXPERIMENT NOTE / 008

EXP 011: Where does a model reviewer earn its keep?
the pre-registration

Every model-reviewer measurement so far used rules a hand-written checker already decides. EXP 011 asks the opposite: on the 18 rules EXP 007's census found no bce 0.3.1 constraint can decide, and which a code change can still break, does nina 0.34.0's reviewer tell violating changes from clean ones better than the frozen rule-filtered off-the-shelf tools, with both held to a false-reject rate of at most 15%? Pre-registered; not yet run.

EXPERIMENT NOTE / 007

EXP 009: Unload leaves no trace:
the results

P1 refuted: 1 of 400 H1 trials diverged from R0 (95% Wilson upper bound 1.40%). P2 supported: 0 silent-inert and 0 false-inactive kernel configurations of 120. All 5 validity gates passed, so the record is informative.

EXPERIMENT NOTE / 006

EXP 009: Unload leaves no trace:
the pre-registration

Does hot load, unload and reconfigure of research components leave the harness indistinguishable from a fresh process, and does every unmet need show up as a named inactive component instead of silently doing nothing? Pre-registered; not yet run.

EXPERIMENT NOTE / 005

EXP 008: Hand over the cache, not the text:
the pre-registration

Two small open-weight models from different families, with different tokenizers. The sender reads a change and the rules once; the receiver must make the EXP 005 gate decision from the sender's KV cache, mapped into its own, without a forward pass over that context. EXP 008 measures whether that matches text re-prefill closely enough, and fast enough, to be worth having, on public baselines only. Pre-registered; not yet run.

EXPERIMENT NOTE / 004

EXP 007: Which of your gate's rules need a model at all?
the pre-registration

For every rule the eight plugins ship that meets the inclusion test (51 items excluded, listed; hunch's 1,011 spec rules sampled 30), how much could a deterministic checker decide on its own, before any model is asked? Pre-registered; measured 7 Oct 2026: the premise is refuted.

EXPERIMENT NOTE / 003

EXP 006: nina reviews the change:
the pre-registration

EXP 005 met nina's spotlight bar, but in 131 of 180 runs its own tool fence kept the reviewer from reading the diff. EXP 006 measures nina 0.34.0's reviewer again on the same 60 changes, with a fence that allows the git forms it uses, and counts a run as having seen the change only when a tool output shows at least one changed line of that change (or, for an added file, the file listed as untracked and at least one of its new lines read). Pre-registered; measured 01 OCT 2026: bar met, nina in the spotlight. Amended 01 Oct 2026, before any counted run.

EXPERIMENT NOTE / 002

Jev as a fast gate:
the pre-registration

From the diff and the rules alone, can Jev turn away architectural drift reliably enough to be a first gate, and hand what it is unsure of to an LLM reviewer? Pre-registered; measured 29 Sep 2026: the claim is refuted.

EXPERIMENT NOTE / 001

System-1 decisions without lock-in:
Jev vs open-weight Laya

Can a local open-weight model answer the same typed decisions, so a hosted model becomes a comparison rather than a dependency? Measured on one laptop, 48 authored rows.

BUILD NOTE / 001

A passing pipeline was not yet
a verified release.

What a missing coverage report taught us about making “done” explainable.

RESEARCH NOTE / 001

Why open the factory floor?

Tools, experiments and a place to show work while it is still becoming something.

The next good question
might be yours.

Bring it to the workbench OPEN
FOR R&D