Metrics
How often the causes the tools report are the causes that were actually planted. Every number is a graded run on a panel whose true structure is known; nothing here is modelled or extrapolated.
321 graded runs across 22 scenario families · panels of 200–12,000 rows × 5–52 columnsAt a glance
A run is fully correct when every planted cause is reported and nothing else is. Precision is the share of reported causes that were planted; recall is the share of planted causes that were reported; F1 combines the two. AUROC ranks every column the tool could have named by the tool's own score and asks how often a planted cause outranks a non-cause.
By tool
The three tools that name causes, graded on the same runs. Time is the wall-clock of the tool call as recorded with each run. 21 of the 321 runs were replayed from a result cache; their times measure the lookup, not the analysis, and are excluded from the time columns.
| Tool | Runs | Fully correct | Precision | Recall | F1 | AUROC | Invented causes | Missed causes | Time (median) | Time (max) |
|---|---|---|---|---|---|---|---|---|---|---|
causal_analysis | 107 | 92% | 0.96 | 0.95 | 0.98 | 0.98 | 9 | 9 of 109 | 29.1 s | 11.8 min |
rank_interventions | 107 | 91% | 0.96 | 0.93 | 0.98 | 0.97 | 5 | 11 of 109 | 3.2 s | 57.9 s |
explain | 107 | 62% | 0.93 | 0.61 | 0.95 | 0.80 | 9 | 45 of 109 | 3.1 s | 1.6 min |
By set
The graded panels fall into five sets. The Traps and the Stress Tests are adversarial: each file was built to make a tool give a confident wrong answer. Pick a set in the By-scenario filter below to see its scenarios one by one.
| Set | Runs | Fully correct | Precision | Recall | F1 | AUROC | Invented causes | Missed causes | Time (median) | Time (max) |
|---|---|---|---|---|---|---|---|---|---|---|
| The Traps files built so that the wrong answer is the tempting one: a reversed cause, a feedback loop, an unmeasured common cause, shared cycles and drift, files with nothing in them, and one real effect fading into noise | 42 | 90% | 0.82 | 0.93 | 1.00 | 0.97 | 3 | 1 of 15 | 1.6 s | 17.9 s |
| Stress Tests one hard property per panel, pushed until something gives: chains, diamonds and fan-outs, long delays, relationships that stop or reverse, a forward-filled sensor, column order, controllers, curved confounding, non-monotone causes, derived columns and wide panels | 180 | 83% | 0.96 | 0.85 | 1.00 | 0.93 | 6 | 26 of 174 | 6.6 s | 11.8 min |
| Baseline scenarios the sample file and its plain variations | 48 | 81% | 1.00 | 0.82 | 0.96 | 0.91 | 0 | 12 of 72 | 3.2 s | 16.7 s |
| Public panels published datasets with public ground truth | 6 | 0% | 0.61 | 0.27 | 0.35 | 0.63 | 13 | 18 of 24 | 75.5 s | 4.7 min |
| Industrial panels plant data with causes planted or documented | 45 | 80% | 0.98 | 0.81 | 0.99 | 0.91 | 1 | 8 of 42 | 5.5 s | 64.4 s |
By scenario
Each family isolates one property a plant dataset can have. Under each tool: the same columns as the By-tool table, on that family's runs. Rows and Columns give the size of the panels in the family. Filter by set, by tool or by scenario; click a column header to sort.
| Scenario family | Set | Tool | Runs | Rows | Columns | Fully correct | Precision | Recall | F1 | AUROC | Invented | Missed | Time (median) | Time (max) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Baseline and variants | The Traps | causal_analysis | 1 | 800 | 5 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 1 | 12.0 s | 12.0 s |
| Baseline and variants | The Traps | rank_interventions | 1 | 800 | 5 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 1 | 1.3 s | 1.3 s |
| Baseline and variants | The Traps | explain | 1 | 800 | 5 | 0% | — | 0.00 | — | 0.50 | 0 | 1 of 1 | 1.7 s | 1.7 s |
| Data with nothing to find | The Traps | causal_analysis | 3 | 800 | 5 | 100% | — | — | — | — | 0 | 0 of 0 | 15.1 s | 17.5 s |
| Data with nothing to find | The Traps | rank_interventions | 3 | 800 | 5 | 100% | — | — | — | — | 0 | 0 of 0 | 0.7 s | 1.1 s |
| Data with nothing to find | The Traps | explain | 3 | 800 | 5 | 100% | — | — | — | — | 0 | 0 of 0 | 1.7 s | 1.7 s |
| Direction of cause | The Traps | causal_analysis | 2 | 800 | 5 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 2 | 16.4 s | 16.7 s |
| Direction of cause | The Traps | rank_interventions | 2 | 800 | 5 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 2 | 0.8 s | 1.0 s |
| Direction of cause | The Traps | explain | 2 | 800 | 5 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 2 | 1.5 s | 1.6 s |
| Known limits | The Traps | causal_analysis | 1 | 800 | 5 | 0% | 0.00 | — | — | — | 1 | 0 of 0 | 17.9 s | 17.9 s |
| Known limits | The Traps | rank_interventions | 1 | 800 | 5 | 0% | 0.00 | — | — | — | 1 | 0 of 0 | 0.8 s | 0.8 s |
| Known limits | The Traps | explain | 1 | 800 | 5 | 0% | 0.00 | — | — | — | 1 | 0 of 0 | 1.6 s | 1.6 s |
| Shared rhythms and drift | The Traps | causal_analysis | 4 | 800 | 5 | 100% | — | — | — | — | 0 | 0 of 0 | 2.4 s | 5.6 s |
| Shared rhythms and drift | The Traps | rank_interventions | 4 | 800 | 5 | 100% | — | — | — | — | 0 | 0 of 0 | 0.9 s | 1.0 s |
| Shared rhythms and drift | The Traps | explain | 4 | 800 | 5 | 100% | — | — | — | — | 0 | 0 of 0 | 1.9 s | 1.9 s |
| Weak signals | The Traps | causal_analysis | 3 | 800 | 5 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 2 | 11.6 s | 11.8 s |
| Weak signals | The Traps | rank_interventions | 3 | 800 | 5 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 2 | 0.9 s | 1.4 s |
| Weak signals | The Traps | explain | 3 | 800 | 5 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 2 | 1.3 s | 1.3 s |
| Closed-loop control | Stress Tests | causal_analysis | 10 | 3,000–6,000 | 5–7 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 10 | 28.8 s | 38.6 s |
| Closed-loop control | Stress Tests | rank_interventions | 10 | 3,000–6,000 | 5–7 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 10 | 3.2 s | 13.7 s |
| Closed-loop control | Stress Tests | explain | 10 | 3,000–6,000 | 5–7 | 80% | 1.00 | 0.83 | 0.96 | 0.92 | 0 | 2 of 10 | 2.9 s | 3.9 s |
| Column order | Stress Tests | causal_analysis | 6 | 5,000 | 9 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 6 | 49.0 s | 52.1 s |
| Column order | Stress Tests | rank_interventions | 6 | 5,000 | 9 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 6 | 5.0 s | 5.2 s |
| Column order | Stress Tests | explain | 6 | 5,000 | 9 | 33% | 1.00 | 0.33 | 1.00 | 0.67 | 0 | 4 of 6 | 4.8 s | 5.3 s |
| Curved (nonlinear) confounding | Stress Tests | causal_analysis | 4 | 2,400 | 5–6 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 4 | 16.8 s | 18.8 s |
| Curved (nonlinear) confounding | Stress Tests | rank_interventions | 4 | 2,400 | 5–6 | 75% | 1.00 | 0.75 | 1.00 | 0.88 | 0 | 1 of 4 | 1.3 s | 1.6 s |
| Curved (nonlinear) confounding | Stress Tests | explain | 4 | 2,400 | 5–6 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 4 | 1.8 s | 2.3 s |
| Data quality | Stress Tests | causal_analysis | 2 | 5,000 | 6 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 2 | 22.1 s | 22.8 s |
| Data quality | Stress Tests | rank_interventions | 2 | 5,000 | 6 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 2 | 2.4 s | 2.4 s |
| Data quality | Stress Tests | explain | 2 | 5,000 | 6 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 2 | 2.3 s | 2.3 s |
| Derived and bookkeeping columns | Stress Tests | causal_analysis | 2 | 2,009–4,000 | 14–16 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 3 | 57.4 s | 63.0 s |
| Derived and bookkeeping columns | Stress Tests | rank_interventions | 2 | 2,009–4,000 | 14–16 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 3 | 16.1 s | 25.2 s |
| Derived and bookkeeping columns | Stress Tests | explain | 2 | 2,009–4,000 | 14–16 | 0% | — | 0.00 | — | 0.50 | 0 | 3 of 3 | 7.9 s | 9.6 s |
| Direct versus indirect causes | Stress Tests | causal_analysis | 7 | 6,000 | 6–8 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 8 | 46.1 s | 56.2 s |
| Direct versus indirect causes | Stress Tests | rank_interventions | 7 | 6,000 | 6–8 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 8 | 4.6 s | 17.4 s |
| Direct versus indirect causes | Stress Tests | explain | 7 | 6,000 | 6–8 | 71% | 1.00 | 0.71 | 1.00 | 0.86 | 0 | 2 of 8 | 4.8 s | 5.3 s |
| Known limits | Stress Tests | causal_analysis | 4 | 2,400–6,000 | 5–6 | 50% | 0.50 | 1.00 | 1.00 | 1.00 | 2 | 0 of 2 | 19.7 s | 26.1 s |
| Known limits | Stress Tests | rank_interventions | 4 | 2,400–6,000 | 5–6 | 50% | 0.50 | 1.00 | 1.00 | 1.00 | 2 | 0 of 2 | 1.8 s | 2.7 s |
| Known limits | Stress Tests | explain | 4 | 2,400–6,000 | 5–6 | 50% | 0.50 | 1.00 | 1.00 | 1.00 | 2 | 0 of 2 | 1.9 s | 2.8 s |
| Long delays | Stress Tests | causal_analysis | 4 | 12,000 | 8 | 75% | 1.00 | 0.75 | 1.00 | 0.88 | 0 | 1 of 4 | 66.1 s | 79.2 s |
| Long delays | Stress Tests | rank_interventions | 4 | 12,000 | 8 | 75% | 1.00 | 0.75 | 1.00 | 0.88 | 0 | 1 of 4 | 4.6 s | 4.7 s |
| Long delays | Stress Tests | explain | 4 | 12,000 | 8 | 75% | 1.00 | 0.75 | 1.00 | 0.88 | 0 | 1 of 4 | 6.6 s | 7.8 s |
| Non-monotone (shape) dependence | Stress Tests | causal_analysis | 5 | 800 | 13 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 5 | 29.9 s | 35.6 s |
| Non-monotone (shape) dependence | Stress Tests | rank_interventions | 5 | 800 | 13 | 80% | 1.00 | 0.80 | 1.00 | 0.90 | 0 | 1 of 5 | 2.3 s | 2.7 s |
| Non-monotone (shape) dependence | Stress Tests | explain | 5 | 800 | 13 | 80% | 1.00 | 0.80 | 1.00 | 0.90 | 0 | 1 of 5 | 5.9 s | 6.4 s |
| Observed confounding | Stress Tests | causal_analysis | 2 | 2,400 | 6 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 2 | 17.2 s | 17.2 s |
| Observed confounding | Stress Tests | rank_interventions | 2 | 2,400 | 6 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 2 | 1.6 s | 1.8 s |
| Observed confounding | Stress Tests | explain | 2 | 2,400 | 6 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 2 | 1.6 s | 1.6 s |
| Regime change | Stress Tests | causal_analysis | 4 | 6,000 | 7 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 3 | 31.5 s | 31.9 s |
| Regime change | Stress Tests | rank_interventions | 4 | 6,000 | 7 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 3 | 3.2 s | 3.4 s |
| Regime change | Stress Tests | explain | 4 | 6,000 | 7 | 50% | 1.00 | 0.33 | 1.00 | 0.67 | 0 | 2 of 3 | 3.2 s | 3.4 s |
| Thin data | Stress Tests | causal_analysis | 1 | 200 | 21 | 100% | — | — | — | — | 0 | 0 of 0 | 56.0 s | 56.0 s |
| Thin data | Stress Tests | rank_interventions | 1 | 200 | 21 | 100% | — | — | — | — | 0 | 0 of 0 | 19.0 s | 19.0 s |
| Thin data | Stress Tests | explain | 1 | 200 | 21 | 100% | — | — | — | — | 0 | 0 of 0 | 19.4 s | 19.4 s |
| Wide panels, few real links | Stress Tests | causal_analysis | 9 | 800–9,600 | 16–51 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 9 | 3.6 min | 11.8 min |
| Wide panels, few real links | Stress Tests | rank_interventions | 9 | 800–9,600 | 16–51 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 9 | 12.0 s | 13.7 s |
| Wide panels, few real links | Stress Tests | explain | 9 | 800–9,600 | 16–51 | 22% | 1.00 | 0.22 | 1.00 | 0.61 | 0 | 7 of 9 | 28.1 s | 79.9 s |
| Baseline and variants | Baseline scenarios | causal_analysis | 10 | 799–800 | 5–10 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 18 | 15.1 s | 16.7 s |
| Baseline and variants | Baseline scenarios | rank_interventions | 10 | 799–800 | 5–10 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 18 | 3.8 s | 7.2 s |
| Baseline and variants | Baseline scenarios | explain | 10 | 799–800 | 5–10 | 60% | 1.00 | 0.61 | 0.94 | 0.81 | 0 | 7 of 18 | 1.7 s | 2.0 s |
| Data with nothing to find | Baseline scenarios | causal_analysis | 1 | 800 | 5 | 100% | — | — | — | — | 0 | 0 of 0 | 14.6 s | 14.6 s |
| Data with nothing to find | Baseline scenarios | rank_interventions | 1 | 800 | 5 | 100% | — | — | — | — | 0 | 0 of 0 | 0.6 s | 0.6 s |
| Data with nothing to find | Baseline scenarios | explain | 1 | 800 | 5 | 100% | — | — | — | — | 0 | 0 of 0 | 1.1 s | 1.1 s |
| Effect sign and chains | Baseline scenarios | causal_analysis | 2 | 800 | 5 | 50% | 1.00 | 0.75 | 0.83 | 0.88 | 0 | 1 of 4 | 8.7 s | 8.9 s |
| Effect sign and chains | Baseline scenarios | rank_interventions | 2 | 800 | 5 | 50% | 1.00 | 0.75 | 0.83 | 0.88 | 0 | 1 of 4 | 2.0 s | 3.3 s |
| Effect sign and chains | Baseline scenarios | explain | 2 | 800 | 5 | 50% | 1.00 | 0.75 | 0.83 | 0.88 | 0 | 1 of 4 | 1.6 s | 1.6 s |
| Levers switched off | Baseline scenarios | causal_analysis | 3 | 800 | 5–6 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 2 | 9.4 s | 11.6 s |
| Levers switched off | Baseline scenarios | rank_interventions | 3 | 800 | 5–6 | 100% | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 0 of 2 | 0.5 s | 0.7 s |
| Levers switched off | Baseline scenarios | explain | 3 | 800 | 5–6 | 33% | — | 0.00 | — | 0.50 | 0 | 2 of 2 | 1.2 s | 1.3 s |
| Public panels | Public panels | causal_analysis | 2 | 960–3,000 | 51–52 | 0% | 0.57 | 0.27 | 0.33 | 0.63 | 6 | 6 of 8 | 3.3 min | 4.7 min |
| Public panels | Public panels | rank_interventions | 2 | 960–3,000 | 51–52 | 0% | 0.67 | 0.27 | 0.38 | 0.62 | 2 | 6 of 8 | 37.3 s | 57.9 s |
| Public panels | Public panels | explain | 2 | 960–3,000 | 51–52 | 0% | 0.58 | 0.27 | 0.34 | 0.64 | 5 | 6 of 8 | 66.3 s | 1.6 min |
| Industrial panels | Industrial panels | causal_analysis | 15 | 3,850 | 10–52 | 93% | 1.00 | 0.93 | 1.00 | 0.96 | 0 | 1 of 14 | 61.1 s | 64.4 s |
| Industrial panels | Industrial panels | rank_interventions | 15 | 3,850 | 10–52 | 93% | 1.00 | 0.93 | 1.00 | 0.96 | 0 | 1 of 14 | 2.7 s | 3.2 s |
| Industrial panels | Industrial panels | explain | 15 | 3,850 | 10–52 | 53% | 0.94 | 0.57 | 0.96 | 0.79 | 1 | 6 of 14 | 5.5 s | 5.9 s |
Public panels
Published industrial datasets with public ground truth, graded the same way. The same scenario and tool filters apply.
| Panel | Set | Tool | Runs | Rows | Columns | Fully correct | Precision | Recall | F1 | AUROC | Invented | Missed | Time (median) | Time (max) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SWaT | Public panels | causal_analysis | 1 | 3,000 | 51 | 0% | 1.00 | 0.33 | 0.50 | 0.75 | 0 | 2 of 3 | 4.7 min | 4.7 min |
| SWaT | Public panels | rank_interventions | 1 | 3,000 | 51 | 0% | 1.00 | 0.33 | 0.50 | 0.67 | 0 | 2 of 3 | 16.7 s | 16.7 s |
| SWaT | Public panels | explain | 1 | 3,000 | 51 | 0% | 1.00 | 0.33 | 0.50 | 0.75 | 0 | 2 of 3 | 1.6 min | 1.6 min |
| Tennessee Eastman | Public panels | causal_analysis | 1 | 960 | 52 | 0% | 0.14 | 0.20 | 0.17 | 0.52 | 6 | 4 of 5 | 2.0 min | 2.0 min |
| Tennessee Eastman | Public panels | rank_interventions | 1 | 960 | 52 | 0% | 0.33 | 0.20 | 0.25 | 0.58 | 2 | 4 of 5 | 57.9 s | 57.9 s |
| Tennessee Eastman | Public panels | explain | 1 | 960 | 52 | 0% | 0.17 | 0.20 | 0.18 | 0.54 | 5 | 4 of 5 | 39.7 s | 39.7 s |
Discovery on published datasets
Full-graph causal discovery against two classical methods, on published datasets with public ground truth and on series generated in the CauseMe style. Same data, same day, same machine. TCPFN appears twice: with the necessary-condition gate on, which is what every tool applies, and with it off, the model alone. Granger and PCMCI carry no gate.
Out of the box: threshold 0.5, no per-dataset tuning
Precision, recall and F1 with each method's normalised score cut at 0.5. AUROC is threshold-free. Edges: reported at the cut / in the ground truth. Time is the method's own run. Bold: best F1 per dataset.
| Dataset | Method | Precision | Recall | F1 at 0.5 | AUROC | Edges (found / true) | Time |
|---|---|---|---|---|---|---|---|
| Tennessee Eastman | TCPFN, gate on | 0.250 | 0.079 | 0.120 | 0.605 | 12 / 38 | 2.0 min |
| Tennessee Eastman | TCPFN, gate off | 0.115 | 0.079 | 0.094 | 0.601 | 26 / 38 | 2.0 min |
| Tennessee Eastman | Granger | 0.016 | 0.921 | 0.031 | 0.603 | 2,241 / 38 | 59.4 s |
| Tennessee Eastman | PCMCI | 0.143 | 0.053 | 0.077 | 0.710 | 14 / 38 | 2.0 min |
| SWaT (water treatment) | TCPFN, gate on | 0.167 | 0.060 | 0.088 | 0.568 | 18 / 50 | 53.7 s |
| SWaT (water treatment) | TCPFN, gate off | 0.096 | 0.100 | 0.098 | 0.568 | 52 / 50 | 54.2 s |
| SWaT (water treatment) | Granger | 0.038 | 0.480 | 0.070 | 0.637 | 631 / 50 | 27.7 s |
| SWaT (water treatment) | PCMCI | 0.500 | 0.060 | 0.107 | 0.500 | 6 / 50 | 43.9 s |
| Causal Rivers | TCPFN, gate on | 0.000 | 0.000 | 0.000 | 0.770 | 3 / 50 | 10.4 min |
| Causal Rivers | TCPFN, gate off | 0.000 | 0.000 | 0.000 | 0.752 | 38 / 50 | 10.4 min |
| Causal Rivers | Granger | 0.006 | 0.880 | 0.013 | 0.705 | 6,817 / 50 | 4.5 min |
| Causal Rivers | PCMCI | 0.000 | 0.000 | 0.000 | 0.581 | 0 / 50 | 6.1 min |
| CAUSRCA (CNC lathe) | TCPFN, gate on | 0.183 | 0.108 | 0.136 | 0.592 | 60 / 102 | 2.8 min |
| CAUSRCA (CNC lathe) | TCPFN, gate off | 0.121 | 0.118 | 0.119 | 0.591 | 99 / 102 | 2.8 min |
| CAUSRCA (CNC lathe) | Granger | 0.020 | 0.275 | 0.038 | 0.536 | 1,366 / 102 | 84.4 s |
| CAUSRCA (CNC lathe) | PCMCI | 0.500 | 0.029 | 0.056 | 0.500 | 6 / 102 | 2.1 min |
| CauseMe-style VAR-5 | TCPFN, gate on | 0.714 | 1.000 | 0.833 | 0.947 | 7 / 5 | 1.3 s |
| CauseMe-style VAR-5 | TCPFN, gate off | 0.714 | 1.000 | 0.833 | 0.933 | 7 / 5 | 1.2 s |
| CauseMe-style VAR-5 | Granger | 0.278 | 1.000 | 0.435 | 1.000 | 18 / 5 | 0.5 s |
| CauseMe-style VAR-5 | PCMCI | 0.000 | 0.000 | 0.000 | 1.000 | 0 / 5 | 0.7 s |
| CauseMe-style VAR-10 | TCPFN, gate on | 0.778 | 0.500 | 0.609 | 0.856 | 18 / 28 | 4.5 s |
| CauseMe-style VAR-10 | TCPFN, gate off | 0.778 | 0.500 | 0.609 | 0.854 | 18 / 28 | 4.5 s |
| CauseMe-style VAR-10 | Granger | 0.325 | 0.929 | 0.481 | 0.858 | 80 / 28 | 2.4 s |
| CauseMe-style VAR-10 | PCMCI | 0.000 | 0.000 | 0.000 | 0.944 | 0 / 28 | 4.0 s |
| CauseMe-style NVAR-5 | TCPFN, gate on | 1.000 | 1.000 | 1.000 | 1.000 | 3 / 3 | 1.2 s |
| CauseMe-style NVAR-5 | TCPFN, gate off | 1.000 | 1.000 | 1.000 | 1.000 | 3 / 3 | 1.0 s |
| CauseMe-style NVAR-5 | Granger | 0.214 | 1.000 | 0.353 | 0.980 | 14 / 3 | 0.5 s |
| CauseMe-style NVAR-5 | PCMCI | 0.000 | 0.000 | 0.000 | 1.000 | 0 / 3 | 0.6 s |
| CauseMe-style NVAR-10 | TCPFN, gate on | 0.786 | 0.478 | 0.595 | 0.883 | 14 / 23 | 4.4 s |
| CauseMe-style NVAR-10 | TCPFN, gate off | 0.786 | 0.478 | 0.595 | 0.880 | 14 / 23 | 4.4 s |
| CauseMe-style NVAR-10 | Granger | 0.265 | 0.957 | 0.415 | 0.919 | 83 / 23 | 2.5 s |
| CauseMe-style NVAR-10 | PCMCI | 0.000 | 0.000 | 0.000 | 0.968 | 0 / 23 | 4.5 s |
| Sachs (biology) | TCPFN, gate on | 0.250 | 0.059 | 0.095 | 0.527 | 4 / 17 | 1.8 min |
| Sachs (biology) | TCPFN, gate off | 0.250 | 0.059 | 0.095 | 0.527 | 4 / 17 | 1.5 min |
| Sachs (biology) | Granger | 0.145 | 0.529 | 0.228 | 0.479 | 62 / 17 | 3.0 s |
| Sachs (biology) | PCMCI | 0.000 | 0.000 | 0.000 | 0.462 | 0 / 17 | 3.4 s |
Best achievable: threshold chosen with the answer key
The best F1 each method reaches when its threshold is tuned on the ground truth, which is not available in a real deployment; the threshold it needed is beside it.
| Dataset | Method | Best F1 | at threshold | AUROC | Edges at best (found / true) |
|---|---|---|---|---|---|
| Tennessee Eastman | TCPFN, gate on | 0.120 | 0.50 | 0.605 | 12 / 38 |
| Tennessee Eastman | TCPFN, gate off | 0.094 | 0.50 | 0.601 | 26 / 38 |
| Tennessee Eastman | Granger | 0.033 | 0.90 | 0.603 | 1,364 / 38 |
| Tennessee Eastman | PCMCI | 0.143 | 0.20 | 0.710 | 60 / 38 |
| SWaT (water treatment) | TCPFN, gate on | 0.118 | 0.30 | 0.568 | 52 / 50 |
| SWaT (water treatment) | TCPFN, gate off | 0.114 | 0.40 | 0.568 | 73 / 50 |
| SWaT (water treatment) | Granger | 0.074 | 0.70 | 0.637 | 546 / 50 |
| SWaT (water treatment) | PCMCI | 0.167 | 0.10 | 0.500 | 34 / 50 |
| Causal Rivers | TCPFN, gate on | 0.053 | 0.20 | 0.770 | 787 / 50 |
| Causal Rivers | TCPFN, gate off | 0.059 | 0.40 | 0.752 | 257 / 50 |
| Causal Rivers | Granger | 0.019 | 0.90 | 0.705 | 3,720 / 50 |
| Causal Rivers | PCMCI | 0.028 | 0.15 | 0.581 | 312 / 50 |
| CAUSRCA (CNC lathe) | TCPFN, gate on | 0.147 | 0.30 | 0.592 | 115 / 102 |
| CAUSRCA (CNC lathe) | TCPFN, gate off | 0.136 | 0.40 | 0.591 | 148 / 102 |
| CAUSRCA (CNC lathe) | Granger | 0.070 | 0.90 | 0.536 | 694 / 102 |
| CAUSRCA (CNC lathe) | PCMCI | 0.176 | 0.15 | 0.500 | 103 / 102 |
| CauseMe-style VAR-5 | TCPFN, gate on | 0.833 | 0.01 | 0.947 | 7 / 5 |
| CauseMe-style VAR-5 | TCPFN, gate off | 0.833 | 0.01 | 0.933 | 7 / 5 |
| CauseMe-style VAR-5 | Granger | 0.714 | 0.90 | 1.000 | 9 / 5 |
| CauseMe-style VAR-5 | PCMCI | 1.000 | 0.10 | 1.000 | 5 / 5 |
| CauseMe-style VAR-10 | TCPFN, gate on | 0.772 | 0.01 | 0.856 | 29 / 28 |
| CauseMe-style VAR-10 | TCPFN, gate off | 0.772 | 0.01 | 0.854 | 29 / 28 |
| CauseMe-style VAR-10 | Granger | 0.562 | 0.90 | 0.858 | 61 / 28 |
| CauseMe-style VAR-10 | PCMCI | 0.943 | 0.10 | 0.944 | 25 / 28 |
| CauseMe-style NVAR-5 | TCPFN, gate on | 1.000 | 0.01 | 1.000 | 3 / 3 |
| CauseMe-style NVAR-5 | TCPFN, gate off | 1.000 | 0.01 | 1.000 | 3 / 3 |
| CauseMe-style NVAR-5 | Granger | 0.667 | 0.90 | 0.980 | 6 / 3 |
| CauseMe-style NVAR-5 | PCMCI | 0.800 | 0.10 | 1.000 | 2 / 3 |
| CauseMe-style NVAR-10 | TCPFN, gate on | 0.800 | 0.01 | 0.883 | 27 / 23 |
| CauseMe-style NVAR-10 | TCPFN, gate off | 0.800 | 0.01 | 0.880 | 27 / 23 |
| CauseMe-style NVAR-10 | Granger | 0.532 | 0.90 | 0.919 | 56 / 23 |
| CauseMe-style NVAR-10 | PCMCI | 0.930 | 0.10 | 0.968 | 20 / 23 |
| Sachs (biology) | TCPFN, gate on | 0.182 | 0.40 | 0.527 | 5 / 17 |
| Sachs (biology) | TCPFN, gate off | 0.182 | 0.40 | 0.527 | 5 / 17 |
| Sachs (biology) | Granger | 0.279 | 0.15 | 0.479 | 105 / 17 |
| Sachs (biology) | PCMCI | 0.272 | 0.01 | 0.462 | 108 / 17 |
Additional datasets
Larger and longer-lag generated series, the water-distribution record and the wind-tunnel record. The two linear long-lag sets plant every edge at delays beyond the lag reach of the gate's independence check, so the gated run refuses them; the ungated row shows what the model alone finds.
| Dataset | Method | Precision | Recall | F1 at 0.5 | Best F1 | AUROC |
|---|---|---|---|---|---|---|
| Synthetic aggregate (50 series) | TCPFN, gate on | 0.670 | 0.115 | 0.182 | — | 0.685 |
| Synthetic aggregate (50 series) | TCPFN, gate off | 0.579 | 0.224 | 0.296 | — | 0.669 |
| Synthetic aggregate (50 series) | Granger | 0.396 | 0.994 | 0.565 | — | 0.602 |
| Synthetic aggregate (50 series) | PCMCI | 0.600 | 0.039 | 0.071 | — | 0.891 |
| BATADAL (water distribution) | TCPFN, gate on | 0.000 | 0.000 | 0.000 | 0.047 | 0.500 |
| BATADAL (water distribution) | TCPFN, gate off | 0.067 | 0.023 | 0.034 | 0.049 | 0.508 |
| BATADAL (water distribution) | Granger | 0.023 | 0.628 | 0.044 | 0.044 | 0.472 |
| BATADAL (water distribution) | PCMCI | 0.500 | 0.023 | 0.044 | 0.044 | 0.473 |
| CauseMe-style VAR-20 | TCPFN, gate on | 0.591 | 0.119 | 0.198 | 0.453 | 0.617 |
| CauseMe-style VAR-20 | TCPFN, gate off | 0.560 | 0.128 | 0.209 | 0.453 | 0.601 |
| CauseMe-style VAR-20 | Granger | 0.287 | 1.000 | 0.446 | 0.446 | 0.509 |
| CauseMe-style VAR-20 | PCMCI | 1.000 | 0.018 | 0.036 | 0.876 | 0.958 |
| CauseMe-style NVAR-20 | TCPFN, gate on | 0.857 | 0.053 | 0.100 | 0.662 | 0.801 |
| CauseMe-style NVAR-20 | TCPFN, gate off | 0.857 | 0.053 | 0.100 | 0.662 | 0.800 |
| CauseMe-style NVAR-20 | Granger | 0.298 | 1.000 | 0.459 | 0.474 | 0.728 |
| CauseMe-style NVAR-20 | PCMCI | 0.000 | 0.000 | 0.000 | 0.910 | 0.958 |
| CauseMe-style Lorenz96-10 | TCPFN, gate on | 0.629 | 0.595 | 0.611 | 0.643 | 0.697 |
| CauseMe-style Lorenz96-10 | TCPFN, gate off | 0.611 | 0.595 | 0.603 | 0.635 | 0.702 |
| CauseMe-style Lorenz96-10 | Granger | 0.411 | 1.000 | 0.583 | 0.587 | 0.580 |
| CauseMe-style Lorenz96-10 | PCMCI | 0.000 | 0.000 | 0.000 | 0.685 | 0.817 |
| CauseMe-style Lorenz96-20 | TCPFN, gate on | 0.507 | 0.387 | 0.439 | 0.504 | 0.693 |
| CauseMe-style Lorenz96-20 | TCPFN, gate off | 0.417 | 0.430 | 0.423 | 0.496 | 0.675 |
| CauseMe-style Lorenz96-20 | Granger | 0.253 | 0.989 | 0.403 | 0.447 | 0.804 |
| CauseMe-style Lorenz96-20 | PCMCI | 0.000 | 0.000 | 0.000 | 0.588 | 0.764 |
| CauseMe-style LongLag-VAR-10 | TCPFN, gate on | 0.000 | 0.000 | 0.000 | 0.423 | 0.538 |
| CauseMe-style LongLag-VAR-10 | TCPFN, gate off | 0.304 | 0.583 | 0.400 | 0.447 | 0.552 |
| CauseMe-style LongLag-VAR-10 | Granger | 0.267 | 1.000 | 0.421 | 0.421 | 0.510 |
| CauseMe-style LongLag-VAR-10 | PCMCI | 0.000 | 0.000 | 0.000 | 0.429 | 0.526 |
| CauseMe-style LongLag-VAR-20 | TCPFN, gate on | 0.000 | 0.000 | 0.000 | 0.414 | 0.475 |
| CauseMe-style LongLag-VAR-20 | TCPFN, gate off | 0.283 | 0.664 | 0.397 | 0.414 | 0.487 |
| CauseMe-style LongLag-VAR-20 | Granger | 0.289 | 1.000 | 0.449 | 0.449 | 0.490 |
| CauseMe-style LongLag-VAR-20 | PCMCI | 0.000 | 0.000 | 0.000 | 0.456 | 0.491 |
| CauseMe-style LongLag-NVAR-10 | TCPFN, gate on | 0.371 | 0.639 | 0.469 | 0.538 | 0.486 |
| CauseMe-style LongLag-NVAR-10 | TCPFN, gate off | 0.371 | 0.639 | 0.469 | 0.533 | 0.481 |
| CauseMe-style LongLag-NVAR-10 | Granger | 0.400 | 1.000 | 0.571 | 0.571 | 0.500 |
| CauseMe-style LongLag-NVAR-10 | PCMCI | 0.000 | 0.000 | 0.000 | 0.567 | 0.445 |
| CauseMe-style LongLag-NVAR-20 | TCPFN, gate on | 0.350 | 0.172 | 0.231 | 0.464 | 0.502 |
| CauseMe-style LongLag-NVAR-20 | TCPFN, gate off | 0.354 | 0.189 | 0.246 | 0.467 | 0.500 |
| CauseMe-style LongLag-NVAR-20 | Granger | 0.321 | 1.000 | 0.486 | 0.486 | 0.500 |
| CauseMe-style LongLag-NVAR-20 | PCMCI | 0.000 | 0.000 | 0.000 | 0.490 | 0.538 |
| Causal Chamber (wind tunnel) | TCPFN, gate on | 0.077 | 0.024 | 0.036 | 0.205 | 0.628 |
| Causal Chamber (wind tunnel) | TCPFN, gate off | 0.056 | 0.048 | 0.051 | 0.203 | 0.622 |
| Causal Chamber (wind tunnel) | Granger | 0.117 | 0.476 | 0.188 | 0.252 | 0.668 |
| Causal Chamber (wind tunnel) | PCMCI | 0.000 | 0.000 | 0.000 | 0.276 | 0.663 |
What each dataset is
| Dataset | Source | Scored on |
|---|---|---|
| Synthetic aggregate (50 series) | generated locally | 10 vars × 500 rows each |
| Tennessee Eastman | published record | 960-row window |
| SWaT (water treatment) | published record | 960-row window |
| Causal Rivers | published record | 960-row window |
| CAUSRCA (CNC lathe) | published record | 960-row window |
| BATADAL (water distribution) | published record | full record |
| CauseMe-style VAR-5 | generated locally | full series |
| CauseMe-style VAR-10 | generated locally | full series |
| CauseMe-style VAR-20 | generated locally | full series |
| CauseMe-style NVAR-5 | generated locally | full series |
| CauseMe-style NVAR-10 | generated locally | full series |
| CauseMe-style NVAR-20 | generated locally | full series |
| CauseMe-style Lorenz96-10 | generated locally | full series |
| CauseMe-style Lorenz96-20 | generated locally | full series |
| CauseMe-style LongLag-VAR-10 | generated locally | full series |
| CauseMe-style LongLag-VAR-20 | generated locally | full series |
| CauseMe-style LongLag-NVAR-10 | generated locally | full series |
| CauseMe-style LongLag-NVAR-20 | generated locally | full series |
| Sachs (biology) | published record | full record |
| Causal Chamber (wind tunnel) | published record | 5,000-row window |
Not shown:
- WADI (water distribution) — ran on a synthetic stand-in; the registered WADI file is not on the run machine.
The judgment-head evaluation from the model's training run is not on this page: it needs the training report, which is not in the repository. Figures published earlier for these datasets were produced before the loaders fetched the genuine records and are not comparable with the rows above.
How to read the numbers
Fully correct means no invented cause and no missed cause on that run.
AUROC is computed per run over the columns the tool examined: planted causes against every other column, scored by the tool's own ranking statistic (a refused or unreported column scores zero). On a panel with a handful of columns it is a coarse number; read it with the counts beside it.
Time is the recorded duration of the tool call on the machine that produced the run, median and maximum over the runs in the row.
Known limits collects the panels built to be unsolvable from their columns alone — an unmeasured common cause, for instance — where the right answer is to report nothing and say why.