ILLUSTRATIVE · NASA / SVS · NOT ADITYA-L1 DATA
Scientific conclusions
A threshold beat the models where it counts.
Machine learning ranks individual minutes better — but it detects no more flares, and it raises about five times as many false alarms. A one-line detector on the SoLEXS count rate remains the operational recommendation.
What is this?
The scientific conclusions of the project, stated at the scale of a claim — including the result that says machine learning did not help.
Why should I trust it?
Every figure on this page is read from the same committed benchmark artifact that the full method tables render from, so the summary cannot drift from the evidence it summarises.
Where can I verify it?
- The proofCurves, calibration, error analysis
- Trace these claimsClaim → artifact → commit
- How this was reachedThe investigation, in order
- Every model that was triedEight cards
Every model, against the simple detector
M/X nowcast · higher is better
- Randomtrivial0.4970.483–0.509
- Majoritytrivial0.5000.500–0.500
- Climatologytrivial0.5000.500–0.500
- Persistencetrivial0.9820.978–0.986
- Threshold (rate)0.9540.940–0.966
- Logistic regression0.9640.953–0.974
- Random forest0.9660.956–0.976
- LightGBM0.9610.949–0.972
ROC-AUC · bars show the 95% confidence interval
Why this result holds
The comparison is decided on the metric that matters operationally — how many real flares are caught, and at what alarm cost — not on the metric that flatters a model.
Decided in advance
seed 20260718
The protocol was frozen before any model was fit.
Extra spectral bands
+0.0033
Δ ROC-AUC from added spectral resolution. A confirmed null.
Held-out test
192,541
Minutes from 2026-01-01 00:00:00+00:00 onward, never seen in training.