Summary
Mean Jaccard (top-5)
0.81
Direction Correct
89.3%
Label Flip Rate
39.3%
Perfect Stability
12/29
Stability by Category
Negation Minimal Pairs (5)
0.7981
Lexical Shortcuts (5)
0.7733
Ambiguity (5)
0.9733
Distribution Style Shift (5)
0.7238
Strong Baselines (5)
0.6729
Edge Cases (4)
0.9428
Per-Input Stability (29 inputs, 5 seeds each)
| ID | Category | Text | Mean Jaccard | Min Jaccard | Top-1 Unanimous |
|---|---|---|---|---|---|
| 1 | Negation Minimal Pairs | I am not entirely unhappy with this result. | 0.5714 | 0.4286 | Yes |
| 2 | Negation Minimal Pairs | This is not good. | 1.0000 | 1.0000 | Yes |
| 3 | Negation Minimal Pairs | This is good. | 1.0000 | 1.0000 | Yes |
| 4 | Negation Minimal Pairs | I would not recommend this to anyone. | 0.6857 | 0.4286 | Yes |
| 5 | Negation Minimal Pairs | Nothing about this experience was disappointing. | 0.7333 | 0.6667 | No |
| 6 | Lexical Shortcuts | The movie was terrible but I loved every minute of it. | 0.5000 | 0.4286 | Yes |
| 7 | Lexical Shortcuts | Excellent packaging, but the product itself is useless. | 0.7667 | 0.6667 | Yes |
| 8 | Lexical Shortcuts | I love how badly this was designed. | 0.8000 | 0.6667 | Yes |
| 9 | Lexical Shortcuts | Fine. | 1.0000 | 1.0000 | Yes |
| 10 | Lexical Shortcuts | Absolutely phenomenal waste of my time. | 0.8000 | 0.6667 | Yes |
| 11 | Ambiguity | It was okay, I guess. | 1.0000 | 1.0000 | Yes |
| 12 | Ambiguity | That is one way to do it. | 0.8667 | 0.6667 | No |
| 13 | Ambiguity | I have seen worse. | 1.0000 | 1.0000 | Yes |
| 14 | Ambiguity | The service was exactly what I expected. | 1.0000 | 1.0000 | No |
| 15 | Ambiguity | It is what it is. | 1.0000 | 1.0000 | Yes |
| 16 | Distribution Style Shift | ngl this slaps fr fr no cap | 1.0000 | 1.0000 | No |
| 17 | Distribution Style Shift | The patient presents with acute exacerbation of chronic symptoms. | 0.5714 | 0.4286 | No |
| 18 | Distribution Style Shift | Revenue increased 12% YoY driven by strong Q4 performance. | 0.5714 | 0.4286 | Yes |
| 19 | Distribution Style Shift | lmaooo this is so bad its good | 0.6762 | 0.4286 | Yes |
| 20 | Distribution Style Shift | Per the attached memo, please advise on next steps. | 0.8000 | 0.6667 | No |
| 21 | Strong Baselines | This is the best product I have ever purchased. | 0.4643 | 0.2500 | No |
| 22 | Strong Baselines | Terrible experience, complete waste of money. | 0.7000 | 0.6667 | Yes |
| 23 | Strong Baselines | I absolutely love everything about this. | 1.0000 | 1.0000 | No |
| 24 | Strong Baselines | This is awful and I regret buying it. | 0.5238 | 0.4286 | No |
| 25 | Strong Baselines | The quality exceeded all my expectations and I am thrilled. | 0.6762 | 0.4286 | Yes |
| 26 | Edge Cases | good good good good good | 1.0000 | 1.0000 | Yes |
| 27 | Edge Cases | The the the the movie was great. | 1.0000 | 1.0000 | Yes |
| 28 | Edge Cases | I think that maybe it could possibly be somewhat decent. | 0.7714 | 0.4286 | No |
| 29 | Edge Cases | Amazing! Horrible! Amazing! Horrible! | 1.0000 | 1.0000 | Yes |
Deletion Faithfulness
Tested
28
Direction Correct
89.3%
Label Flips
11/28
Mean |Delta|
0.3846
Direction-Incorrect Counterexamples (3 of 28)
In these cases, LIME predicted that removing a token would shift the score in one direction, but the actual effect went the opposite way. This happens when the model has redundant evidence and no single token is individually necessary for the prediction.
#5“Nothing about this experience was disappointing.”
Direction IncorrectLIME assigned a negative weight to "experience" (predicting removal would increase positive score), but removing it slightly decreased the positive score. The model's confidence was so high that removing any single token had negligible effect.
#22“Terrible experience, complete waste of money.”
Direction IncorrectLIME assigned a negative weight to "Terrible" (predicting removal would increase positive score), but removing it had virtually zero effect. The model's confidence was so high that removing any single token had negligible effect.
#26“good good good good good”
Direction IncorrectLIME assigned a positive weight to "good" (predicting removal would decrease positive score), but removing it had virtually zero effect. Redundant evidence means no single token is individually necessary.
Methodology
Protocol pre-registered before data collection. 29 inputs across 6 categories, each run through LIME with 5 different random seeds (42, 123, 456, 789, 1024) producing 145 total LIME attribution sets.
Stability measured via pairwise Jaccard similarity on top-5 tokens across all C(5,2)=10 seed pairs per input. Faithfulness measured via single-token deletion: remove the highest-weighted token and measure whether the confidence shift matches the predicted direction.
Thresholds: Jaccard ≥ 0.6 (pass), direction correctness ≥ 70% (pass). Full protocol and failure analysis available in the repository.