XAI Forensics

Pre-registered audit results: 29 inputs, 5 seeds, 145 LIME runs

Summary

Mean Jaccard (top-5)

0.81

Direction Correct

89.3%

Label Flip Rate

39.3%

Perfect Stability

12/29

Stability by Category

Negation Minimal Pairs (5)

0.7981

Lexical Shortcuts (5)

0.7733

Ambiguity (5)

0.9733

Distribution Style Shift (5)

0.7238

Strong Baselines (5)

0.6729

Edge Cases (4)

0.9428

Per-Input Stability (29 inputs, 5 seeds each)

IDCategoryTextMean JaccardMin JaccardTop-1 Unanimous
1Negation Minimal PairsI am not entirely unhappy with this result.0.57140.4286Yes
2Negation Minimal PairsThis is not good.1.00001.0000Yes
3Negation Minimal PairsThis is good.1.00001.0000Yes
4Negation Minimal PairsI would not recommend this to anyone.0.68570.4286Yes
5Negation Minimal PairsNothing about this experience was disappointing.0.73330.6667No
6Lexical ShortcutsThe movie was terrible but I loved every minute of it.0.50000.4286Yes
7Lexical ShortcutsExcellent packaging, but the product itself is useless.0.76670.6667Yes
8Lexical ShortcutsI love how badly this was designed.0.80000.6667Yes
9Lexical ShortcutsFine.1.00001.0000Yes
10Lexical ShortcutsAbsolutely phenomenal waste of my time.0.80000.6667Yes
11AmbiguityIt was okay, I guess.1.00001.0000Yes
12AmbiguityThat is one way to do it.0.86670.6667No
13AmbiguityI have seen worse.1.00001.0000Yes
14AmbiguityThe service was exactly what I expected.1.00001.0000No
15AmbiguityIt is what it is.1.00001.0000Yes
16Distribution Style Shiftngl this slaps fr fr no cap1.00001.0000No
17Distribution Style ShiftThe patient presents with acute exacerbation of chronic symptoms.0.57140.4286No
18Distribution Style ShiftRevenue increased 12% YoY driven by strong Q4 performance.0.57140.4286Yes
19Distribution Style Shiftlmaooo this is so bad its good0.67620.4286Yes
20Distribution Style ShiftPer the attached memo, please advise on next steps.0.80000.6667No
21Strong BaselinesThis is the best product I have ever purchased.0.46430.2500No
22Strong BaselinesTerrible experience, complete waste of money.0.70000.6667Yes
23Strong BaselinesI absolutely love everything about this.1.00001.0000No
24Strong BaselinesThis is awful and I regret buying it.0.52380.4286No
25Strong BaselinesThe quality exceeded all my expectations and I am thrilled.0.67620.4286Yes
26Edge Casesgood good good good good1.00001.0000Yes
27Edge CasesThe the the the movie was great.1.00001.0000Yes
28Edge CasesI think that maybe it could possibly be somewhat decent.0.77140.4286No
29Edge CasesAmazing! Horrible! Amazing! Horrible!1.00001.0000Yes

Deletion Faithfulness

Tested

28

Direction Correct

89.3%

Label Flips

11/28

Mean |Delta|

0.3846

Direction-Incorrect Counterexamples (3 of 28)

In these cases, LIME predicted that removing a token would shift the score in one direction, but the actual effect went the opposite way. This happens when the model has redundant evidence and no single token is individually necessary for the prediction.

#5Nothing about this experience was disappointing.

Direction Incorrect
Removed Tokenexperience
LIME Weight-0.1825
Actual Delta-0.002447
Label ChangeNo flip

LIME assigned a negative weight to "experience" (predicting removal would increase positive score), but removing it slightly decreased the positive score. The model's confidence was so high that removing any single token had negligible effect.

#22Terrible experience, complete waste of money.

Direction Incorrect
Removed TokenTerrible
LIME Weight-0.3175
Actual Delta-0.000002
Label ChangeNo flip

LIME assigned a negative weight to "Terrible" (predicting removal would increase positive score), but removing it had virtually zero effect. The model's confidence was so high that removing any single token had negligible effect.

#26good good good good good

Direction Incorrect
Removed Tokengood
LIME Weight+0.0210
Actual Delta+0.000001
Label ChangeNo flip

LIME assigned a positive weight to "good" (predicting removal would decrease positive score), but removing it had virtually zero effect. Redundant evidence means no single token is individually necessary.

Methodology

Protocol pre-registered before data collection. 29 inputs across 6 categories, each run through LIME with 5 different random seeds (42, 123, 456, 789, 1024) producing 145 total LIME attribution sets.

Stability measured via pairwise Jaccard similarity on top-5 tokens across all C(5,2)=10 seed pairs per input. Faithfulness measured via single-token deletion: remove the highest-weighted token and measure whether the confidence shift matches the predicted direction.

Thresholds: Jaccard ≥ 0.6 (pass), direction correctness ≥ 70% (pass). Full protocol and failure analysis available in the repository.