VINDEX3

model.vindex/ — A VINDEX3 CONTAINER

A model, stored as a directory you can read. One file — index.json — speaks for the whole container; system_graph.json carries the judged meaning — components, logical objects, and each layer's declared operator; the segments carry the bytes, one per logical object. Every part named, findable, checkable.

OPEN THE CONTAINER →

VINDEX3 · 3.0 CANDIDATE SPECIFICATION

THE MODEL IS THE DATABASE

A self-describing, executable, queryable model container: the same copy can be run, questioned, checked — and changed, with proof. Nothing re-exported for each use, nothing thrown away.

A model you can run.
A computation you can inspect.
A claim you can test.

An AI model is billions of learned numbers, and today's formats keep those numbers perfectly — as storage. What they do not keep is everything else the release meant: which parts are which, what may consume them, which precisions are still the same model, what was ever proven about any of it. VINDEX3 keeps the numbers and the meaning — every part named, every representation catalogued, every claim checkable — for the life of the artifact.

A modern model release is not a weights file. It is a system.

ONE RELEASE — TWO INTERPRETATIONS

the same checkpoint, byte-identically preserved either way

A WEIGHTS FILEA DATABASE
  • Addresses model semantics — knows what they mean
  • Each layer's operator declared — surfaces follow the program
  • Representations present, selected, authoritative
  • Run it, query it, verify it — the same bytes

the same checkpoint, byte-identically preserved either way — as a weights file: addresses stored tensors — knows where they are, one precision, chosen once at conversion, meaning lives in filename conventions. As a database: addresses model semantics — knows what they mean, each layer's operator declared — surfaces follow the program, representations present, selected, authoritative, run it, query it, verify it — the same bytes.

If that claim is true, you should be able to ask the file itself.

ONE QUERY, STRAIGHT AT THE WEIGHTS

WALK "the capital of France" TOP 3

No forward pass, and no separate index — the answer is read from the stored gate rows themselves, layer by layer. WALK and DESCRIBE are the browse surface the ABI itself specifies. Try it, live, in the Explorer →

A worked shape, not a recorded run. The browse surface ships today as an analysis-only profile; expert-region browse parity is still open — the Record keeps score.

THE EXECUTION IS A RECORD.

If a container knows how to execute, the execution can identify its writes. Those writes can become a record.

ONE EXECUTION, LEFT AS EVIDENCE · GEMMA 3 4B

“The capital of France is”

Follow the recorded readout for token #9079 — “ Paris” in the capture study.

L0RECORDED PROBABILITY · 0–100%L33
Token rank
1
Lens probability
69.74%
Carrier norm · L2
40683.05
Applied-write norm · L2
3227.98

Rank 1 at layer 24. Almost 100% at layer 26. Back to 80% at the final layer. A readout can strengthen, then weaken. The record lets you return to each write.

READ THE PROVENANCE & SCOPE

Recorded 2026-09-20 · production CPU · final prompt position 5 · head-v1 lens. 408 carrier writes, 204 readouts, 1,446 events. Hardware model is not recorded.

Stored BF16; the production image also pins Q8 requantisation for 103 operands. A representation label alone does not describe the arithmetic.

Prompt text, model name and token spelling come from the capture study. The JSONL stores token IDs and model identity. Probability is exp(recorded log p); this is a lens readout, not generated output or a causal claim.

Run: gemma3-4b-france-lens
File SHA-256: 12cdf4d9d06f5675233d71556c2158702cff88dd89efe42af72ab9dfb345dbd4

In Observatory, choose “Open Paris recording”, or open the downloaded JSONL. This exhibit reads a checked excerpt; Observatory validates and replays the original record. Neither runs inference.

The execution left a record.
No model needs to be running to inspect it again.

OBSERVE, ATTRIBUTE, INTERVENE →

Where does such a file come from? It is compiled — once.

EXTRACT ONCE — THE WHOLE BRIDGE, PERFORMED

A checkpoint compiles down. Then it is no longer needed.

a checkpoint — what you download today

config.json
model-00001-of-00004.safetensors
model-00002-of-00004.safetensors
model-00003-of-00004.safetensors
model-00004-of-00004.safetensors
tokenizer.json

the checkpoint may now be deleted — execution must not change

inventory ↓

plan ↓

graph ↓

encode ↓

verify ↓

model.vindex/ — written, then proven

system_graph.json
segments/ — one per logical object
tokenizer.json + capability snapshot
index.json — the root, written last

verified — byte-faithful to its source

A checkpoint — config.json and safetensors shards — is inventoried, judged, formed into a graph, encoded in write order with index.json last, and verified against its source. Then the checkpoint may be deleted: execution must not change. That is the whole bridge, and it is crossed once.

106 tokens per second, from one container, on one laptop — and the answer, provably unchanged.

gpt-oss-20b · one M3 Max · measured 2026-08-20 · same greedy ids on every arm — accounted on the Record →

REPRESENT · EVIDENCE-DIRECTED COMPILATION

Precision is something a model earns.

Freeze the behaviour to preserve. Compile a candidate. Measure the composed model against a reference. Let the evidence decide which physical form is admissible.

  1. IDENTIFY CANDIDATE
  2. COMPILE
  3. SEALED TOKEN BANK
  4. MEASURE
  5. INGEST EVIDENCE
  6. ADJUDICATE
  7. SEARCH · REFUSE · PROMOTE

RECORDED MAP · 256 POSITIONS

L20–25 Q8 · L26 Q6

KL p99
2.489e-3
Limit
1.000e-3

REFUSED

Each member passed alone. Their composition exceeds the KL budget.

RECORDED MAP · 8,192 POSITIONS

L24–26 → Q8_0

KL p99
4.153e-4
Limit
1.000e-3

PASSED THE FROZEN CONTRACT

L0–23 stay BF16. The whole model passes all six criteria, not just KL.

Kimi-Linear-48B-A3B-Instruct · 2026-08-30 · kimi-logit-v3 · teacher-forced evaluation. Diagnostic and authority scales are labelled separately. Dated precision-topology evidence; these are not results from the new AUTO-REP campaign.

The current measurement tool establishes whether evidence is admissible. Acceptance needs a separately declared gate. AUTO-REP refuses to run an unarmed plan; a search interface does not mean a campaign has passed.

THE CLAIM CAN BE TESTED.

WHAT KIND OF THING DO WE KNOW?

  1. DECLARED

    The graph says this operator exists.

    A description of the program, before it runs.

  2. OBSERVED

    This carrier was written.

    A value captured at a named execution boundary.

  3. ATTRIBUTED

    This head contributes under this readout.

    Descriptive support depends on the reader and normalization contract.

  4. INTERVENED

    Changing this head changed the continuation.

    Requires a recorded manipulation and a controlled comparison.

  5. TESTED

    This effect held within the frozen experiment.

    A scoped result. Its controls, subject and decision rule travel with the claim.

These are different claims, each needing its own evidence. Observation does not establish attribution; attribution does not establish a counterfactual. Even an intervention needs controls before it supports a causal conclusion.

THE EVIDENCE CONTRACT →

One family. Distinct jobs.

VINDEX3
The artifact. A model is an executable database.
LARQL
The engine. Query and operate that database.
OBSERVATORY
The instrument. See what its computation did.
REPRESENT
The compilation programme. Test changes to its physical form against declared behaviour.
HAUSE
The language. Make the structure, the evidence and the refusals legible.

Or skip the reading and put your hands on it — the surfaces answer from the same knowledge the chapters teach, and the CLI runs it all on your own machine.