Results & benchmarks

Results with their receipts attached.

A result is not a mood, a roadmap, or a source file. This page is a bounded index of measurements, evaluations, simulations, and execution receipts, with the workload, source, evidence type, and limitation kept visible.

Measurement rule

No numeric result is invented here. A benchmark belongs with its hardware, revision, workload, method, baseline, and limitation. When that record is not public, this page names the gap instead of filling it with a promise.

Evidence index

Different evidence answers different questions.

These entries link to the surface that can be inspected. They are deliberately not ranked against one another.

RECEIPT · NATIVE EXECUTION

Neural Foundry

Question: what local execution path was requested, resolved, accepted, observed, and persisted?

Record: bounded execution receipts, source contracts, tests, and checkpoint/integrity surfaces where available.

Limitation: a source receipt does not prove target-hardware acceptance or promotion.

Read the execution boundary →
EVALUATION · PUBLIC HARNESS

Open World Model Harness

Question: how does a system behave over time when the evaluator, host world, model output, and policy are kept separate?

Record: reproducible evaluation and replay artifacts within the public harness boundary.

Limitation: harness evidence is not a production deployment claim.

Inspect the evaluation shelf →
SIMULATION · DEVELOPMENT EVIDENCE

OMNI-Q

Question: can a simulation-oriented loop route objectives, verify work, recover from failure, and retain evidence?

Record: mock and simulation results identified as simulation.

Limitation: simulation does not establish trained-model capability, hardware performance, or universal behavior.

See the Showcase entry →
SHOWCASE · CREATION

VANTA and Cartographer

Question: what does the creation stack currently expose?

Record: product boundaries, authoring/runtime relationships, and implementation status attached to observable entries.

Limitation: a concept, product direction, or case study is not automatically a measured feature.

Inspect creation evidence →
RESEARCH · PUBLICATION

SSRN and technical writing

Question: what public argument, method, or institutional analysis is available to read?

Record: author, title, DOI or source URL, publication date, and overview page when one exists.

Limitation: a publication is not a software benchmark or a company performance guarantee.

Open the research library →
SOURCE · RELEASE RECORD

Public repository receipts

Question: what public source revision was reviewed for the current directory?

Record: repository identity, reviewed revision, release note, and non-claims in the public portfolio manifest.

Limitation: a commit hash is provenance, not proof that every local or hardware path ran.

Read source receipts →
Required context

Every serious result needs a reconstruction path.

Workload Model, dataset, scenario, fixture, or publication scope.
Environment Hardware, runtime, driver, operating system, or simulator version when relevant.
Revision Source commit, binary identity, request, plan, or document record.
Method Baseline, comparison, measurement procedure, and retained receipt.
Limitations What the evidence does not establish, including any unverified deployment or promotion claim.
Evidence boundary

Results are one layer of the system map. Use the Research Software index for investigations, the Showcase for observable work, and the Open Canopy contract for public release boundaries.