Predictions with honest uncertainty

Same upside class, a third of the risk.

P1-Chaos is a specialised machine learning engine for classification on structured tabular data. It combines iterative feedback architecture, spectral feature extraction, and calibrated ensemble methods to produce probabilistic predictions at the precision that operational decisions require.

Chaotic Hierarchical AutoML with Oscillating Signals
At a glance
Survival C-index (GBSG2)
0.681

Best Uno IPCW-C on scikit-survival’s canonical set

Lower peak drawdown

BTC-USD walk-forward 2015–2025, net of 10 bps — 0.28 max drawdown vs 0.89 buy-and-hold.

Net Sharpe after costs
1.33

vs 0.82 buy-and-hold on the same asset and window — ~22% annualised vol vs ~67%.

Survival benchmark rank
Leads

Best Uno C-index on GBSG2 (0.681) — canonical scikit-survival set; published DeepHit: 0.675.

A sharper prediction is not a better decision if the band that carries it lies — promising 80% and delivering 46%. The problem is not beating every leaderboard. It is shipping uncertainty you can size a position, a schedule, or a commitment against.

Most models end at a point prediction and treat every regime, every asset, and every confidence level as if the error structure were stable. It is not. Chaos rebuilds the decision on three layers — conformal guarantees that mean what they say, a causal regime engine that sets the risk budget to match conditions, and calibrated duration intervals that stay honest as markets drift. Where a raw forecast leaves you guessing, Chaos begins

Principia One
01

Calibrated uncertainty

Split-conformal and adaptive (ACI) layers. Worst native under-coverage across 28 cells: −34%; wrapped: −3.6%. Cross-checked against MAPIE on UCI.

02

Regime-conditioned risk

Four causal regimes set conviction and risk budget. BTC walk-forward drawdown 0.27 vs 0.89 buy & hold comes from here.

03

Regime duration

Persistence as a calibrated survival problem. 80% nominal intervals near 80% empirical coverage — planning, not a silent position gate.

The floor on a log-log line.

Bitcoin’s calendar-year lows, aged from the 2009-01-03 genesis and fit as P ≈ C · t^β. Same valuation anchor the Chaos-Crypto power-law overlay uses — a floor that has risen with network age, not a price oracle. Overlay remains OFF by default; see the robustness report.

[ BTC ANNUAL LOWS · LOG–LOG POWER LAW ]Valuation anchor · not a forecast
// CHAOS / REGIMES

Four market regimes, one risk policy

Chaos labels every bar into one of four causal market regimes and sets conviction and vol budget to match. On BTC the edge is regime-conditioned risk control — not duration timing and not a direction oracle.

01

Strong trend

Persistent directional move with high trend strength. Chaos amplifies conviction and sizes up within the vol budget.

Policy
Amplify momentum (gain ×1.4)
Share of days
21.4%
Avg exposure
0.54
Chaos Sharpe
2.60
02

Risk on

Constructive, orderly tape. Standard momentum conviction with a measured, vol-targeted position.

Policy
Standard momentum (gain ×1.0)
Share of days
37.7%
Avg exposure
0.21
Chaos Sharpe
1.00
03

Mean revert

Choppy, range-bound conditions where momentum fails. Exposure stays small; production policy fades short-term moves.

Policy
Fade momentum (gain ×−0.6)
Share of days
13.7%
Avg exposure
0.09
Chaos Sharpe
04

Risk off

Stress and elevated volatility. Conviction pulls toward flat — capital protection first.

Policy
De-risk toward flat (gain ×0.3, vol ×1.5)
Share of days
27.1%
Avg exposure
0.02
Chaos Sharpe
0.71

Regime policy ablations (BTC-USD)

Pre-registered variants on the same causal walk-forward — production, naked control, return-oriented regime tweaks, and the frozen trend-gated power-law overlay. No parameter sweep.

VariantNet SharpeTotal returnAnn. returnMax DDAvg |pos|Δ vs prod.
Production book
Vol-targeted regime-conditioned position — the committed marketing harness.
1.30615.0×30%0.2660.212
Same alpha, no regime layer
Fixed risk_on label, vol target off, trend filter off — Panel B naked control.
0.9829.0×23%0.2810.255-0.324
Strong-trend vol budget 70%
Pre-registered: 50% ann. vol budget everywhere except 70% in strong_trend.
1.25121.4×34%0.3240.242-0.055
Mean-revert flat (no fade)
Pre-registered: mean_revert gain 0 instead of −0.6 — stop fighting chop.
1.36616.9×31%0.2420.201+0.060
Risk-off floor in uptrend
Pre-registered: min 10% long exposure in risk_off when 200d trend_filter > 0.3 — participate in bulls mis-labeled as stress.
1.27814.2×29%0.2620.217-0.028
Production + trend-gated power-law (full history)
Frozen BTC overlay: λ=0.3, scale=1.5, trend gate <0.3. Full 2015–2025 window — not comparable to crypto_powerlaw.md Sharpe (that harness starts at 40% burn-in).
Full 2015–2025 window; overlay hurts early-era return vs production — see powerlaw_gated_oos for the committed powerlaw harness window. (2015-07-20 → 2025-12-29)
1.1505.7×18%0.2220.174-0.156
Power-law overlay (powerlaw OOS window)
Same frozen overlay; walk-forward from max(40% history, 260 bars) — matches synthetic_paths / crypto_powerlaw.md protocol.
OOS window matches crypto_powerlaw.md (40% burn-in, min 260 bars). (2019-05-26 → 2025-12-29)
0.9472.3×14%0.1670.180-0.359

The BTC walk-forward book is sized by regime only. Regime-duration forecasts are a separate calibrated-interval product (see Duration); they do not gate crypto positions.

Calibrated regime duration

A separate survival layer for interval honesty on multi-asset panels — not used to size the BTC trading book.

Nominal interval
80%
Empirical coverage
80.5%
Honest read

Duration beats a memoryless null on several series and keeps ACI intervals near 80% nominal — but gating BTC positions on duration does not beat a shuffled control (see regime_duration_crypto_value.md). Stated plainly: calibration for planning, not crypto Sharpe.

BTC-USD · primary80% interval

Risk off

~8 more days in regime

[0, 16] days

Primary crypto series — planning interval only, not a position gate.

Gold80% interval

Strong trend

~15 more days in regime

[6, 24] days
SPY80% interval

Risk on

~7 more days in regime

[0, 14] days
US10Y80% interval

Risk off

~10 more days in regime

[1, 19] days
VIX80% interval

Mean revert

~11 more days in regime

[0, 30] days
// CHAOS / REGIME DETECTION

The regime engine, run over real history

The production regime engine — causal, 252-day lookback, no look-ahead — labels every week of real price history into one of four regimes. The bands below are what it saw at the time, not a fit after the fact: the 2018 and 2022 crypto winters, the 2008 financial crisis and the 2020 COVID crash all read risk-off; the late-2024 melt-up reads strong-trend.

Strong trendAmplify (gain ×1.4)Risk onStandard momentum (×1.0)Mean revertFade (gain −0.6)Risk offDe-risk toward flat
BitcoinBTC-USD · crypto · 603 weeks · risk-off 51% of weeks
Current regime
Risk off
Spell age
20d
Expected more
~8d
80% interval
[0, 16]d
as of 2025-12-30 · planning only
20162017201820192020202120222023202420252026
EthereumETH-USD · crypto · 454 weeks · risk-off 54% of weeks
20192020202120222023202420252026
S&P 500 (SPY)SPY · equity · 1020 weeks · risk-off 40% of weeks
2008200920102011201220132014201520162017201820192020202120222023202420252026
GoldGC=F · safe-haven · 1020 weeks · risk-off 47% of weeks
2008200920102011201220132014201520162017201820192020202120222023202420252026

Detection generalizes to any series — shown here on crypto, an equity index, and gold. The trading edge is validated on crypto-class assets only; labeling SPY is a detection demonstration, not a claim the signal trades equities profitably.

get_regime_labels (Mason + ADX + Hurst/Kalman/HMM fusion), weekly-dominant regime; display spans under 3 weeks absorbed for legibility. — Source: p1-chaos benchmarks/regime_timeline.py → reports/regime_timeline.json

At a glance

Most models end at a point prediction and treat every regime, every asset, and every confidence level as if the error structure were stable. It is not. Chaos rebuilds the decision on three layers — conformal guarantees that mean what they say, a causal regime engine that sets the risk budget to match conditions, and calibrated duration intervals that stay honest as markets drift. Where a raw forecast leaves you guessing, Chaos begins

[ CHAOS · PREDICTION VS REALITY ]Held-out backtest
[ CHAOS VS BUY & HOLD · WALK-FORWARD ]Net of 10 bps costs

Chaos vs the field — validated on canonical public benchmarks

Chaos is optimised for calibrated probabilities and risk-adjusted control. On recognized public benchmarks its genuine, externally-verified edges are (1) survival / time-to-event on canonical medical datasets and (2) crypto risk-adjusted economics. On general tabular classification and univariate forecasting it is competitive but does not beat tuned standard baselines (CatBoost, seasonal methods) — stated plainly.

Survival C-index — GBSG2 (canonical)leads
0.681vs 0.675 (DeepHit (published))

Uno IPCW C-index, ≥ the published deep-survival SOTA on the standard GBSG2 dataset

Survival calibration — GBSG2 (IBS)leads
0.165vs 0.170 (published Cox/DeepSurv)

lower integrated Brier than the published reference range (0.170–0.184)

Crypto net Sharpe (after costs)leads
0.65vs 0.36 (Buy & Hold)
6-asset arena, Friedman rank

Friedman rank 1.8 of 6; ~3× lower drawdown at 1/3 the volatility

Forecast interval coverage (80% nominal)calibrated
87%vs (conformal)

split-conformal + ACI 80% intervals on the real financial panel

[ ARENA SCORECARD · 4 ARENAS ]
ArenaChaosBest competitorVerdict
Survival (canonical public datasets)
Uno IPCW C-index ↑ / integrated Brier ↓
Chaos-Survival
GBSG2 C-index 0.681 · IBS 0.165
CoxPH / RSF / GBSA / DeepHit
Cox 0.660 · RSF 0.680 · DeepHit (pub) 0.675
leads

On scikit-survival's canonical GBSG2 (the standard survival benchmark), Chaos-Survival has the best Uno C-index and a lower integrated Brier than the published deep-survival reference range — a genuine external win. Competitive (close 2nd) on WHAS500 (0.773) and Veterans (0.680). This is the real calibration story.

Crypto
Net Sharpe after 10bps costs (higher = better)
Chaos-Crypto
0.65
Buy & Hold
0.36
leads

Best risk-adjusted return and ~3–4× lower drawdown. Beats EWMA-Momentum, HAR-RV and a flat book (Holm p<0.05). Honest: nobody has a daily-direction edge (AUC ≈ 0.51) — the edge is risk control, not calling price.

Forecasting (M4, official OWA)
OWA (lower = better)
Chaos-RegimeBlend / Chaos-Global
M4-Hourly OWA 0.46 (blend); M4-Daily 0.91 (global, best)
Seasonal methods / AutoTheta
varies by frequency
competitive

Beats the statistical baselines on several M4 frequencies and the global model wins M4-Daily, but no uniform win — a seasonal-naive baseline is best on Hourly. Differentiator is calibrated intervals, not point accuracy.

Tabular classification (OpenML-CC18)
ROC-AUC ↑ / log-loss / Brier / ECE ↓
Chaos-Stacked
AUC 0.846 · Brier 0.116 · ECE 0.056
CatBoost
AUC 0.856 · Brier 0.111 · ECE 0.052
trails

On the canonical OpenML-CC18 suite, CatBoost leads on every metric including calibration (ECE/Brier). Chaos's calibration edge on our internal arena did NOT generalize here — shown honestly. A follow-up landscape run that adds the open foundation model TabPFN v2 found TabPFN now tops both CatBoost and Chaos on accuracy AND calibration (see chaos-landscape.json). Recommended use: wrap the best base model — CatBoost or TabPFN — with the Chaos conformal + regime layer rather than replace it.

Numbers are reproduced from clean runs on canonical open-source datasets. Where Chaos trails a standard baseline, it is shown as 'competitive' or 'trails', never inflated. Source: p1-chaos benchmarks + canonical open-source evals (scikit-survival datasets, OpenML-CC18, M4).

Receipts

Each promise on this page is measured as a delta over the unwrapped practice a practitioner would actually use — at the same 80% nominal the copy quotes. Every number regenerates from a committed harness (p1-chaos benchmarks/wrapper_value.py); wins and losses land as they fall.

[ NATIVE → CHAOS · AT 80% NOMINAL ]

The wrapper's calibration cost (20% of training data) is charged in every Panel A cell, not hidden. Panel B is an ablation of our own signal, not an external comparison. Full tables, including the cells we lose, live in the committed report.

Source: p1-chaos benchmarks/wrapper_value.py → reports/wrapper_value.md