TCPFN v2.1.1.50.3
MCP API REST API Performance Metrics

Metrics

How often the causes the tools report are the causes that were actually planted. Every number is a graded run on a panel whose true structure is known; nothing here is modelled or extrapolated.

321 graded runs across 22 scenario families · panels of 200–12,000 rows × 5–52 columns

At a glance

A run is fully correct when every planted cause is reported and nothing else is. Precision is the share of reported causes that were planted; recall is the share of planted causes that were reported; F1 combines the two. AUROC ranks every column the tool could have named by the tool's own score and asks how often a planted cause outranks a non-cause.

321
graded runs
22
scenario families
82%
runs fully correct
0.97
mean F1
0.92
mean AUROC
200–12,000
rows per panel
5–52
columns per panel

By tool

The three tools that name causes, graded on the same runs. Time is the wall-clock of the tool call as recorded with each run. 21 of the 321 runs were replayed from a result cache; their times measure the lookup, not the analysis, and are excluded from the time columns.

ToolRuns Fully correctPrecision RecallF1 AUROCInvented causes Missed causesTime (median) Time (max)
causal_analysis10792%0.960.950.980.9899 of 10929.1 s11.8 min
rank_interventions10791%0.960.930.980.97511 of 1093.2 s57.9 s
explain10762%0.930.610.950.80945 of 1093.1 s1.6 min

By set

The graded panels fall into five sets. The Traps and the Stress Tests are adversarial: each file was built to make a tool give a confident wrong answer. Pick a set in the By-scenario filter below to see its scenarios one by one.

SetRuns Fully correctPrecision RecallF1 AUROCInvented causes Missed causesTime (median) Time (max)
The Traps

files built so that the wrong answer is the tempting one: a reversed cause, a feedback loop, an unmeasured common cause, shared cycles and drift, files with nothing in them, and one real effect fading into noise

4290%0.820.931.000.9731 of 151.6 s17.9 s
Stress Tests

one hard property per panel, pushed until something gives: chains, diamonds and fan-outs, long delays, relationships that stop or reverse, a forward-filled sensor, column order, controllers, curved confounding, non-monotone causes, derived columns and wide panels

18083%0.960.851.000.93626 of 1746.6 s11.8 min
Baseline scenarios

the sample file and its plain variations

4881%1.000.820.960.91012 of 723.2 s16.7 s
Public panels

published datasets with public ground truth

60%0.610.270.350.631318 of 2475.5 s4.7 min
Industrial panels

plant data with causes planted or documented

4580%0.980.810.990.9118 of 425.5 s64.4 s

By scenario

Each family isolates one property a plant dataset can have. Under each tool: the same columns as the By-tool table, on that family's runs. Rows and Columns give the size of the panels in the family. Filter by set, by tool or by scenario; click a column header to sort.

Scenario familySetToolRunsRowsColumnsFully correctPrecisionRecallF1AUROCInventedMissedTime (median)Time (max)
Baseline and variantsThe Trapscausal_analysis18005100%1.001.001.001.0000 of 112.0 s12.0 s
Baseline and variantsThe Trapsrank_interventions18005100%1.001.001.001.0000 of 11.3 s1.3 s
Baseline and variantsThe Trapsexplain180050%0.000.5001 of 11.7 s1.7 s
Data with nothing to findThe Trapscausal_analysis38005100%00 of 015.1 s17.5 s
Data with nothing to findThe Trapsrank_interventions38005100%00 of 00.7 s1.1 s
Data with nothing to findThe Trapsexplain38005100%00 of 01.7 s1.7 s
Direction of causeThe Trapscausal_analysis28005100%1.001.001.001.0000 of 216.4 s16.7 s
Direction of causeThe Trapsrank_interventions28005100%1.001.001.001.0000 of 20.8 s1.0 s
Direction of causeThe Trapsexplain28005100%1.001.001.001.0000 of 21.5 s1.6 s
Known limitsThe Trapscausal_analysis180050%0.0010 of 017.9 s17.9 s
Known limitsThe Trapsrank_interventions180050%0.0010 of 00.8 s0.8 s
Known limitsThe Trapsexplain180050%0.0010 of 01.6 s1.6 s
Shared rhythms and driftThe Trapscausal_analysis48005100%00 of 02.4 s5.6 s
Shared rhythms and driftThe Trapsrank_interventions48005100%00 of 00.9 s1.0 s
Shared rhythms and driftThe Trapsexplain48005100%00 of 01.9 s1.9 s
Weak signalsThe Trapscausal_analysis38005100%1.001.001.001.0000 of 211.6 s11.8 s
Weak signalsThe Trapsrank_interventions38005100%1.001.001.001.0000 of 20.9 s1.4 s
Weak signalsThe Trapsexplain38005100%1.001.001.001.0000 of 21.3 s1.3 s
Closed-loop controlStress Testscausal_analysis103,000–6,0005–7100%1.001.001.001.0000 of 1028.8 s38.6 s
Closed-loop controlStress Testsrank_interventions103,000–6,0005–7100%1.001.001.001.0000 of 103.2 s13.7 s
Closed-loop controlStress Testsexplain103,000–6,0005–780%1.000.830.960.9202 of 102.9 s3.9 s
Column orderStress Testscausal_analysis65,0009100%1.001.001.001.0000 of 649.0 s52.1 s
Column orderStress Testsrank_interventions65,0009100%1.001.001.001.0000 of 65.0 s5.2 s
Column orderStress Testsexplain65,000933%1.000.331.000.6704 of 64.8 s5.3 s
Curved (nonlinear) confoundingStress Testscausal_analysis42,4005–6100%1.001.001.001.0000 of 416.8 s18.8 s
Curved (nonlinear) confoundingStress Testsrank_interventions42,4005–675%1.000.751.000.8801 of 41.3 s1.6 s
Curved (nonlinear) confoundingStress Testsexplain42,4005–6100%1.001.001.001.0000 of 41.8 s2.3 s
Data qualityStress Testscausal_analysis25,0006100%1.001.001.001.0000 of 222.1 s22.8 s
Data qualityStress Testsrank_interventions25,0006100%1.001.001.001.0000 of 22.4 s2.4 s
Data qualityStress Testsexplain25,0006100%1.001.001.001.0000 of 22.3 s2.3 s
Derived and bookkeeping columnsStress Testscausal_analysis22,009–4,00014–16100%1.001.001.001.0000 of 357.4 s63.0 s
Derived and bookkeeping columnsStress Testsrank_interventions22,009–4,00014–16100%1.001.001.001.0000 of 316.1 s25.2 s
Derived and bookkeeping columnsStress Testsexplain22,009–4,00014–160%0.000.5003 of 37.9 s9.6 s
Direct versus indirect causesStress Testscausal_analysis76,0006–8100%1.001.001.001.0000 of 846.1 s56.2 s
Direct versus indirect causesStress Testsrank_interventions76,0006–8100%1.001.001.001.0000 of 84.6 s17.4 s
Direct versus indirect causesStress Testsexplain76,0006–871%1.000.711.000.8602 of 84.8 s5.3 s
Known limitsStress Testscausal_analysis42,400–6,0005–650%0.501.001.001.0020 of 219.7 s26.1 s
Known limitsStress Testsrank_interventions42,400–6,0005–650%0.501.001.001.0020 of 21.8 s2.7 s
Known limitsStress Testsexplain42,400–6,0005–650%0.501.001.001.0020 of 21.9 s2.8 s
Long delaysStress Testscausal_analysis412,000875%1.000.751.000.8801 of 466.1 s79.2 s
Long delaysStress Testsrank_interventions412,000875%1.000.751.000.8801 of 44.6 s4.7 s
Long delaysStress Testsexplain412,000875%1.000.751.000.8801 of 46.6 s7.8 s
Non-monotone (shape) dependenceStress Testscausal_analysis580013100%1.001.001.001.0000 of 529.9 s35.6 s
Non-monotone (shape) dependenceStress Testsrank_interventions58001380%1.000.801.000.9001 of 52.3 s2.7 s
Non-monotone (shape) dependenceStress Testsexplain58001380%1.000.801.000.9001 of 55.9 s6.4 s
Observed confoundingStress Testscausal_analysis22,4006100%1.001.001.001.0000 of 217.2 s17.2 s
Observed confoundingStress Testsrank_interventions22,4006100%1.001.001.001.0000 of 21.6 s1.8 s
Observed confoundingStress Testsexplain22,4006100%1.001.001.001.0000 of 21.6 s1.6 s
Regime changeStress Testscausal_analysis46,0007100%1.001.001.001.0000 of 331.5 s31.9 s
Regime changeStress Testsrank_interventions46,0007100%1.001.001.001.0000 of 33.2 s3.4 s
Regime changeStress Testsexplain46,000750%1.000.331.000.6702 of 33.2 s3.4 s
Thin dataStress Testscausal_analysis120021100%00 of 056.0 s56.0 s
Thin dataStress Testsrank_interventions120021100%00 of 019.0 s19.0 s
Thin dataStress Testsexplain120021100%00 of 019.4 s19.4 s
Wide panels, few real linksStress Testscausal_analysis9800–9,60016–51100%1.001.001.001.0000 of 93.6 min11.8 min
Wide panels, few real linksStress Testsrank_interventions9800–9,60016–51100%1.001.001.001.0000 of 912.0 s13.7 s
Wide panels, few real linksStress Testsexplain9800–9,60016–5122%1.000.221.000.6107 of 928.1 s79.9 s
Baseline and variantsBaseline scenarioscausal_analysis10799–8005–10100%1.001.001.001.0000 of 1815.1 s16.7 s
Baseline and variantsBaseline scenariosrank_interventions10799–8005–10100%1.001.001.001.0000 of 183.8 s7.2 s
Baseline and variantsBaseline scenariosexplain10799–8005–1060%1.000.610.940.8107 of 181.7 s2.0 s
Data with nothing to findBaseline scenarioscausal_analysis18005100%00 of 014.6 s14.6 s
Data with nothing to findBaseline scenariosrank_interventions18005100%00 of 00.6 s0.6 s
Data with nothing to findBaseline scenariosexplain18005100%00 of 01.1 s1.1 s
Effect sign and chainsBaseline scenarioscausal_analysis2800550%1.000.750.830.8801 of 48.7 s8.9 s
Effect sign and chainsBaseline scenariosrank_interventions2800550%1.000.750.830.8801 of 42.0 s3.3 s
Effect sign and chainsBaseline scenariosexplain2800550%1.000.750.830.8801 of 41.6 s1.6 s
Levers switched offBaseline scenarioscausal_analysis38005–6100%1.001.001.001.0000 of 29.4 s11.6 s
Levers switched offBaseline scenariosrank_interventions38005–6100%1.001.001.001.0000 of 20.5 s0.7 s
Levers switched offBaseline scenariosexplain38005–633%0.000.5002 of 21.2 s1.3 s
Public panelsPublic panelscausal_analysis2960–3,00051–520%0.570.270.330.6366 of 83.3 min4.7 min
Public panelsPublic panelsrank_interventions2960–3,00051–520%0.670.270.380.6226 of 837.3 s57.9 s
Public panelsPublic panelsexplain2960–3,00051–520%0.580.270.340.6456 of 866.3 s1.6 min
Industrial panelsIndustrial panelscausal_analysis153,85010–5293%1.000.931.000.9601 of 1461.1 s64.4 s
Industrial panelsIndustrial panelsrank_interventions153,85010–5293%1.000.931.000.9601 of 142.7 s3.2 s
Industrial panelsIndustrial panelsexplain153,85010–5253%0.940.570.960.7916 of 145.5 s5.9 s

Public panels

Published industrial datasets with public ground truth, graded the same way. The same scenario and tool filters apply.

PanelSetToolRunsRowsColumnsFully correctPrecisionRecallF1AUROCInventedMissedTime (median)Time (max)
SWaTPublic panelscausal_analysis13,000510%1.000.330.500.7502 of 34.7 min4.7 min
SWaTPublic panelsrank_interventions13,000510%1.000.330.500.6702 of 316.7 s16.7 s
SWaTPublic panelsexplain13,000510%1.000.330.500.7502 of 31.6 min1.6 min
Tennessee EastmanPublic panelscausal_analysis1960520%0.140.200.170.5264 of 52.0 min2.0 min
Tennessee EastmanPublic panelsrank_interventions1960520%0.330.200.250.5824 of 557.9 s57.9 s
Tennessee EastmanPublic panelsexplain1960520%0.170.200.180.5454 of 539.7 s39.7 s

Discovery on published datasets

Full-graph causal discovery against two classical methods, on published datasets with public ground truth and on series generated in the CauseMe style. Same data, same day, same machine. TCPFN appears twice: with the necessary-condition gate on, which is what every tool applies, and with it off, the model alone. Granger and PCMCI carry no gate.

final_v2.1.safetensors · NVIDIA GeForce RTX 5080 · run 2026-09-18 · threshold 0.5 · max lag 3 · hybrid prune on

Out of the box: threshold 0.5, no per-dataset tuning

Precision, recall and F1 with each method's normalised score cut at 0.5. AUROC is threshold-free. Edges: reported at the cut / in the ground truth. Time is the method's own run. Bold: best F1 per dataset.

DatasetMethodPrecisionRecallF1 at 0.5AUROCEdges (found / true)Time
Tennessee EastmanTCPFN, gate on0.2500.0790.1200.60512 / 382.0 min
Tennessee EastmanTCPFN, gate off0.1150.0790.0940.60126 / 382.0 min
Tennessee EastmanGranger0.0160.9210.0310.6032,241 / 3859.4 s
Tennessee EastmanPCMCI0.1430.0530.0770.71014 / 382.0 min
SWaT (water treatment)TCPFN, gate on0.1670.0600.0880.56818 / 5053.7 s
SWaT (water treatment)TCPFN, gate off0.0960.1000.0980.56852 / 5054.2 s
SWaT (water treatment)Granger0.0380.4800.0700.637631 / 5027.7 s
SWaT (water treatment)PCMCI0.5000.0600.1070.5006 / 5043.9 s
Causal RiversTCPFN, gate on0.0000.0000.0000.7703 / 5010.4 min
Causal RiversTCPFN, gate off0.0000.0000.0000.75238 / 5010.4 min
Causal RiversGranger0.0060.8800.0130.7056,817 / 504.5 min
Causal RiversPCMCI0.0000.0000.0000.5810 / 506.1 min
CAUSRCA (CNC lathe)TCPFN, gate on0.1830.1080.1360.59260 / 1022.8 min
CAUSRCA (CNC lathe)TCPFN, gate off0.1210.1180.1190.59199 / 1022.8 min
CAUSRCA (CNC lathe)Granger0.0200.2750.0380.5361,366 / 10284.4 s
CAUSRCA (CNC lathe)PCMCI0.5000.0290.0560.5006 / 1022.1 min
CauseMe-style VAR-5TCPFN, gate on0.7141.0000.8330.9477 / 51.3 s
CauseMe-style VAR-5TCPFN, gate off0.7141.0000.8330.9337 / 51.2 s
CauseMe-style VAR-5Granger0.2781.0000.4351.00018 / 50.5 s
CauseMe-style VAR-5PCMCI0.0000.0000.0001.0000 / 50.7 s
CauseMe-style VAR-10TCPFN, gate on0.7780.5000.6090.85618 / 284.5 s
CauseMe-style VAR-10TCPFN, gate off0.7780.5000.6090.85418 / 284.5 s
CauseMe-style VAR-10Granger0.3250.9290.4810.85880 / 282.4 s
CauseMe-style VAR-10PCMCI0.0000.0000.0000.9440 / 284.0 s
CauseMe-style NVAR-5TCPFN, gate on1.0001.0001.0001.0003 / 31.2 s
CauseMe-style NVAR-5TCPFN, gate off1.0001.0001.0001.0003 / 31.0 s
CauseMe-style NVAR-5Granger0.2141.0000.3530.98014 / 30.5 s
CauseMe-style NVAR-5PCMCI0.0000.0000.0001.0000 / 30.6 s
CauseMe-style NVAR-10TCPFN, gate on0.7860.4780.5950.88314 / 234.4 s
CauseMe-style NVAR-10TCPFN, gate off0.7860.4780.5950.88014 / 234.4 s
CauseMe-style NVAR-10Granger0.2650.9570.4150.91983 / 232.5 s
CauseMe-style NVAR-10PCMCI0.0000.0000.0000.9680 / 234.5 s
Sachs (biology)TCPFN, gate on0.2500.0590.0950.5274 / 171.8 min
Sachs (biology)TCPFN, gate off0.2500.0590.0950.5274 / 171.5 min
Sachs (biology)Granger0.1450.5290.2280.47962 / 173.0 s
Sachs (biology)PCMCI0.0000.0000.0000.4620 / 173.4 s

Best achievable: threshold chosen with the answer key

The best F1 each method reaches when its threshold is tuned on the ground truth, which is not available in a real deployment; the threshold it needed is beside it.

DatasetMethodBest F1at thresholdAUROCEdges at best (found / true)
Tennessee EastmanTCPFN, gate on0.1200.500.60512 / 38
Tennessee EastmanTCPFN, gate off0.0940.500.60126 / 38
Tennessee EastmanGranger0.0330.900.6031,364 / 38
Tennessee EastmanPCMCI0.1430.200.71060 / 38
SWaT (water treatment)TCPFN, gate on0.1180.300.56852 / 50
SWaT (water treatment)TCPFN, gate off0.1140.400.56873 / 50
SWaT (water treatment)Granger0.0740.700.637546 / 50
SWaT (water treatment)PCMCI0.1670.100.50034 / 50
Causal RiversTCPFN, gate on0.0530.200.770787 / 50
Causal RiversTCPFN, gate off0.0590.400.752257 / 50
Causal RiversGranger0.0190.900.7053,720 / 50
Causal RiversPCMCI0.0280.150.581312 / 50
CAUSRCA (CNC lathe)TCPFN, gate on0.1470.300.592115 / 102
CAUSRCA (CNC lathe)TCPFN, gate off0.1360.400.591148 / 102
CAUSRCA (CNC lathe)Granger0.0700.900.536694 / 102
CAUSRCA (CNC lathe)PCMCI0.1760.150.500103 / 102
CauseMe-style VAR-5TCPFN, gate on0.8330.010.9477 / 5
CauseMe-style VAR-5TCPFN, gate off0.8330.010.9337 / 5
CauseMe-style VAR-5Granger0.7140.901.0009 / 5
CauseMe-style VAR-5PCMCI1.0000.101.0005 / 5
CauseMe-style VAR-10TCPFN, gate on0.7720.010.85629 / 28
CauseMe-style VAR-10TCPFN, gate off0.7720.010.85429 / 28
CauseMe-style VAR-10Granger0.5620.900.85861 / 28
CauseMe-style VAR-10PCMCI0.9430.100.94425 / 28
CauseMe-style NVAR-5TCPFN, gate on1.0000.011.0003 / 3
CauseMe-style NVAR-5TCPFN, gate off1.0000.011.0003 / 3
CauseMe-style NVAR-5Granger0.6670.900.9806 / 3
CauseMe-style NVAR-5PCMCI0.8000.101.0002 / 3
CauseMe-style NVAR-10TCPFN, gate on0.8000.010.88327 / 23
CauseMe-style NVAR-10TCPFN, gate off0.8000.010.88027 / 23
CauseMe-style NVAR-10Granger0.5320.900.91956 / 23
CauseMe-style NVAR-10PCMCI0.9300.100.96820 / 23
Sachs (biology)TCPFN, gate on0.1820.400.5275 / 17
Sachs (biology)TCPFN, gate off0.1820.400.5275 / 17
Sachs (biology)Granger0.2790.150.479105 / 17
Sachs (biology)PCMCI0.2720.010.462108 / 17

Additional datasets

Larger and longer-lag generated series, the water-distribution record and the wind-tunnel record. The two linear long-lag sets plant every edge at delays beyond the lag reach of the gate's independence check, so the gated run refuses them; the ungated row shows what the model alone finds.

DatasetMethodPrecisionRecallF1 at 0.5Best F1AUROC
Synthetic aggregate (50 series)TCPFN, gate on0.6700.1150.1820.685
Synthetic aggregate (50 series)TCPFN, gate off0.5790.2240.2960.669
Synthetic aggregate (50 series)Granger0.3960.9940.5650.602
Synthetic aggregate (50 series)PCMCI0.6000.0390.0710.891
BATADAL (water distribution)TCPFN, gate on0.0000.0000.0000.0470.500
BATADAL (water distribution)TCPFN, gate off0.0670.0230.0340.0490.508
BATADAL (water distribution)Granger0.0230.6280.0440.0440.472
BATADAL (water distribution)PCMCI0.5000.0230.0440.0440.473
CauseMe-style VAR-20TCPFN, gate on0.5910.1190.1980.4530.617
CauseMe-style VAR-20TCPFN, gate off0.5600.1280.2090.4530.601
CauseMe-style VAR-20Granger0.2871.0000.4460.4460.509
CauseMe-style VAR-20PCMCI1.0000.0180.0360.8760.958
CauseMe-style NVAR-20TCPFN, gate on0.8570.0530.1000.6620.801
CauseMe-style NVAR-20TCPFN, gate off0.8570.0530.1000.6620.800
CauseMe-style NVAR-20Granger0.2981.0000.4590.4740.728
CauseMe-style NVAR-20PCMCI0.0000.0000.0000.9100.958
CauseMe-style Lorenz96-10TCPFN, gate on0.6290.5950.6110.6430.697
CauseMe-style Lorenz96-10TCPFN, gate off0.6110.5950.6030.6350.702
CauseMe-style Lorenz96-10Granger0.4111.0000.5830.5870.580
CauseMe-style Lorenz96-10PCMCI0.0000.0000.0000.6850.817
CauseMe-style Lorenz96-20TCPFN, gate on0.5070.3870.4390.5040.693
CauseMe-style Lorenz96-20TCPFN, gate off0.4170.4300.4230.4960.675
CauseMe-style Lorenz96-20Granger0.2530.9890.4030.4470.804
CauseMe-style Lorenz96-20PCMCI0.0000.0000.0000.5880.764
CauseMe-style LongLag-VAR-10TCPFN, gate on0.0000.0000.0000.4230.538
CauseMe-style LongLag-VAR-10TCPFN, gate off0.3040.5830.4000.4470.552
CauseMe-style LongLag-VAR-10Granger0.2671.0000.4210.4210.510
CauseMe-style LongLag-VAR-10PCMCI0.0000.0000.0000.4290.526
CauseMe-style LongLag-VAR-20TCPFN, gate on0.0000.0000.0000.4140.475
CauseMe-style LongLag-VAR-20TCPFN, gate off0.2830.6640.3970.4140.487
CauseMe-style LongLag-VAR-20Granger0.2891.0000.4490.4490.490
CauseMe-style LongLag-VAR-20PCMCI0.0000.0000.0000.4560.491
CauseMe-style LongLag-NVAR-10TCPFN, gate on0.3710.6390.4690.5380.486
CauseMe-style LongLag-NVAR-10TCPFN, gate off0.3710.6390.4690.5330.481
CauseMe-style LongLag-NVAR-10Granger0.4001.0000.5710.5710.500
CauseMe-style LongLag-NVAR-10PCMCI0.0000.0000.0000.5670.445
CauseMe-style LongLag-NVAR-20TCPFN, gate on0.3500.1720.2310.4640.502
CauseMe-style LongLag-NVAR-20TCPFN, gate off0.3540.1890.2460.4670.500
CauseMe-style LongLag-NVAR-20Granger0.3211.0000.4860.4860.500
CauseMe-style LongLag-NVAR-20PCMCI0.0000.0000.0000.4900.538
Causal Chamber (wind tunnel)TCPFN, gate on0.0770.0240.0360.2050.628
Causal Chamber (wind tunnel)TCPFN, gate off0.0560.0480.0510.2030.622
Causal Chamber (wind tunnel)Granger0.1170.4760.1880.2520.668
Causal Chamber (wind tunnel)PCMCI0.0000.0000.0000.2760.663

What each dataset is

Dataset SourceScored on
Synthetic aggregate (50 series)generated locally10 vars × 500 rows each
Tennessee Eastmanpublished record960-row window
SWaT (water treatment)published record960-row window
Causal Riverspublished record960-row window
CAUSRCA (CNC lathe)published record960-row window
BATADAL (water distribution)published recordfull record
CauseMe-style VAR-5generated locallyfull series
CauseMe-style VAR-10generated locallyfull series
CauseMe-style VAR-20generated locallyfull series
CauseMe-style NVAR-5generated locallyfull series
CauseMe-style NVAR-10generated locallyfull series
CauseMe-style NVAR-20generated locallyfull series
CauseMe-style Lorenz96-10generated locallyfull series
CauseMe-style Lorenz96-20generated locallyfull series
CauseMe-style LongLag-VAR-10generated locallyfull series
CauseMe-style LongLag-VAR-20generated locallyfull series
CauseMe-style LongLag-NVAR-10generated locallyfull series
CauseMe-style LongLag-NVAR-20generated locallyfull series
Sachs (biology)published recordfull record
Causal Chamber (wind tunnel)published record5,000-row window

Not shown:

  • WADI (water distribution) — ran on a synthetic stand-in; the registered WADI file is not on the run machine.

The judgment-head evaluation from the model's training run is not on this page: it needs the training report, which is not in the repository. Figures published earlier for these datasets were produced before the loaders fetched the genuine records and are not comparable with the rows above.

How to read the numbers

Planted truth. Each graded panel is generated with known causes, so a reported cause is either planted (a hit), invented (not planted), or missed (planted but not reported). Some panels also mark causes as acceptable but not required — an indirect cause, or a duplicate sensor — and those count as neither hits nor inventions.
Fully correct means no invented cause and no missed cause on that run.
AUROC is computed per run over the columns the tool examined: planted causes against every other column, scored by the tool's own ranking statistic (a refused or unreported column scores zero). On a panel with a handful of columns it is a coarse number; read it with the counts beside it.
Time is the recorded duration of the tool call on the machine that produced the run, median and maximum over the runs in the row.
Known limits collects the panels built to be unsolvable from their columns alone — an unmeasured common cause, for instance — where the right answer is to report nothing and say why.