How often are the agents right?
Every day we save what the agents said, wait, and then check it against what actually happened. No cherry-picking — every call counts.
Live Track Record
5-day horizonEach daily signal is recorded and scored 5 trading days later: a bullish call is “correct” if the price is higher, a bearish call if lower. This is the live record — it grows every day and cannot be edited. Collecting since 7/13/2026.
These four numbers are the deterministic advisor’s alone — the engine behind today’s signals. The retired editorial baseline is kept in the record for honesty and reported separately below; the two are never added together. Across every stream, 245 calls have been evaluated and 65 are pending.
Which engine produced the calls
Every row in the record is labeled with the stack that produced it, and the two streams are never blended into one number. “Deterministic advisor” is the live quant engine (measured candles, ADX, volatility, calibrated confidence) that generates today’s signals; “editorial baseline” is the fallback it replaced, kept here because its history is part of the honest record.
| Signal engine | Recorded | Evaluated | Pending | Accuracy | 95% range | Avg move |
|---|---|---|---|---|---|---|
| Deterministic advisor (the quant engine) | 64 | 55 | 9 | 44% | 31–57% | -0.603% |
| Editorial baseline (pre-engine fallback) | 246 | 190 | 56 | 45% | 38–52% | +0.181% |
Accuracy by market
No accuracy claim is made below 30 evaluated signals — until then this page reports collection progress only. For the multi-year simulation of the same logic, see the admin Backtest Lab. Past accuracy does not guarantee future results.
Does the stated confidence mean anything?
calibration error 26ptsThe receipt behind every confidence number: signals are grouped by what we said (stated confidence) and scored by what happened (measured hit rate). Perfect honesty means the two columns match. Calibration is reported per engine — a calibrated stream blended with an uncalibrated one would describe neither. This table is machine-readable at /api/track-record/reliability and cannot be edited.
Deterministic advisor (the quant engine)
55 evaluated signals; calibration error 7 points. 9 pending.
| We said | Signals | Measured hit rate | 95% range | Status |
|---|---|---|---|---|
| 45–50% | 19 | — | — | collecting 19 of 30 |
| 50–55% | 36 | 47% | 32–63% | measured |
Editorial baseline (pre-engine fallback)
190 evaluated signals; calibration error 31 points. 56 pending.
| We said | Signals | Measured hit rate | 95% range | Status |
|---|---|---|---|---|
| 55–60% | 2 | — | — | collecting 2 of 30 |
| 60–65% | 15 | — | — | collecting 15 of 30 |
| 65–70% | 32 | 53% | 36–69% | measured |
| 70–75% | 35 | 54% | 38–70% | measured |
| 75–80% | 34 | 35% | 22–52% | measured |
| 80–85% | 69 | 38% | 27–50% | measured |
| 85–90% | 3 | — | — | collecting 3 of 30 |
Reliability = measured hit rate per stated-confidence bucket over the live, append-only signal record, reported per signal stream: 'advisor_live' is the deterministic quant engine, 'seed_baseline' the editorial fallback it replaced. Buckets under 30 evaluated signals are still collecting and make no claim. Past accuracy does not guarantee future results. 'brier' splits forecast error via isotonic recalibration (CORP): miscalibration is the fixable part, discrimination is genuine ordering skill, uncertainty is the irreducible base-rate variance; identity brier = miscalibration - discrimination + uncertainty.
Can people actually read this?
measured, not assertedWe claim the product explains itself to beginners. Onboarding ends with three questions about how to read a card — what a model score is, what a stop means, and what “follow with paper money” does. These are the answers, including the wrong ones. A rate only counts as a claim at 50 answers per mode.
No one has taken the check yet. It appears at the end of onboarding.
Comprehension is a claim only at N >= 50 per mode; below that this is collection progress.