OAM Chess Lab
Beta Research · Empirical Milestone

OAM and Stockfish are measuring different things

An initial empirical comparison between OAM's Moment detector and conventional Stockfish engine criticality across five historical chess games. Read this as a beta research finding, not a marketing claim.

0 / 14
OAM Moments on the engine's primary-critical ply
12 / 14
OAM Moments at < 50 cp engine swing
r = 0.077
Correlation · OE_SCORE ↔ engine swing

5 games · 14 OAM Moments · Stockfish 15.1 · depth 16 · repository OAM package unchanged

In the tested corpus, 0 of 14 OAM Moments coincided with the engine's primary critical ply, and 12 of 14 occurred at positions with less than 50 centipawns of engine evaluation swing. The correlation between OAM's OE_SCORE and engine evaluation swing was r = 0.077.

These results do not show that OAM is better than Stockfish. Stockfish and OAM answer different questions, and Stockfish is a mature chess-engine system while OAM is an experimental measurement framework.

What the result does show is sufficiently interesting to investigate further: OAM's current Moment detector does not appear, in this initial test, to simply reproduce conventional engine criticality.

The next question is whether players and coaches find this different observable useful.

This is beta research. The experiment is continuing.

Inspect the experiment

If you are an IM, coach, chess researcher, or technically minded player, you can read the full report and the raw dataset directly. The production OAM package (`/app/backend/oam/`) is used unchanged; the experimental runner does not modify it.

  • REPORT.md — full A–I methodology, common-coordinate dataset, A/B/C/D classification, Evergreen deep-dive, limitations
  • dataset.json — every OAM Moment and every per-ply engine metric (114 kB)
  • oam_vs_engine.py — experimental runner (Python · reproducible with Stockfish 15.1)

Falsifiable proposition tested

"OAM's current Moment detector is substantially reducible to conventional engine criticality."

Verdict on this corpus: WEAKENED

Read descriptively, not competitively. See § I of the report for corpus-size and configuration limits.

OAM-100 corpusOAM-only analysisStockfish-only analysisMeasurement comparisonMeasurement divergence
Beta Research · OAM-100 · Measurement Comparison

Same trajectories. Different measurements.

What did the experiment ask?

Do OAM and Stockfish measure the same function of the chess trajectory?

What did it find?

Not supported, under this experiment.

The same 100 human-game trajectories were analysed independently by the frozen OAM implementation and Stockfish 15.1 at depth 16. Their measurements showed weak event correspondence, near-zero magnitude association, poor reciprocal coverage, and very low deterministic reconstructability. OAM-selected transitions also tended to occur at substantially lower Stockfish evaluation movement than the overall non-mate trajectory.

OAM-100 experiment comparing OAM and Stockfish measurements across six correspondence tests, showing measurement divergence under the experiment.
OAM-100 visual abstract · same trajectories · different measurements
1 · Event correspondence
Weak
OAM ∩ SF top-3 / game = 1.09%
2 · Magnitude correspondence
Weak
Spearman OE vs ΔSF = −0.019
3 · Distribution correspondence
Moderate
Mean ΔSF at OAM Moments: 12.61 cp
Mean ΔSF over all non-mate plies: 23.64 cp
Ratio: 0.53
4 · Reciprocal correspondence
Weak
P(OAM | SF ≥ 100 cp) ≈ 1%
5 · Reconstructability
Weak
Best deterministic F1 (either direction) = 0.017
6 · Null-model resistance
Not significant
at α = 0.05, two-sided
Observed mean: 12.61 cp
Null mean: 25.12 ± 5.56 cp
p = 0.0515
OAM-onlyunder this experiment
  • 174 / 183 non-mate OAM Moments occurred at transitions where Stockfish registered less than 50 cp of evaluation movement under the tested configuration ( 95.1% ).
  • OE_SCORE variation was essentially orthogonal to ΔSF (Spearman ρ ≈ 0).
  • FTS / RDS / RLS receive-response structure was not recovered by Stockfish.
Stockfish-onlyunder this experiment
  • ≥ 50 cp events: 837 / 846 (99%) not flagged by OAM.
  • ≥ 100 cp events: 303 / 306 (99%) not flagged by OAM.
  • ≥ 200 cp events: 73 / 73 (100%) not flagged by OAM.
  • ≥ 300 cp events: 33 / 33 (100%) not flagged by OAM.
  • Per-game Top-1 engine swing: 0 / 100 flagged by OAM.

This establishes measurement divergence under the experiment. It does not establish OAM superiority, consciousness, general validity, or the nature of the phenomenon OAM detects.

Experimental context
Corpus
100 human games
Legal plies
7,813
OAM Moments
190
Aligned non-mate
183
Engine
Stockfish 15.1
Depth
16
Threads
1
Hash
32 MB
MultiPV
1
Perspective
White POV
Comparison
OAM vs engine evaluation swing
Null model
10,000 cluster-preserving within-game permutations
Random seed
20260215
Reproducibility
Byte-identical across two invocations
Inspect the comparison

The frozen Gate-4 artefacts are inspectable here. The corpus, OAM output and Stockfish output referenced by the comparison remain unchanged.

Try it on a game

Stockfish tells you the best move · OAM tells you the moment your game changed · research continues
● Public Model · OAM Engine v1.0oamchess.com