Two papers, one question:
can you trust the highlight?
A model reads a Python file and says "Author 7 wrote this". A highlighter then colours the parts of the code that supposedly made it say so. This page explains, from the ground up, how both of the thesis papers test whether those colours tell the truth. It starts with everything the two papers share, then splits into the first paper and the latest one.
The model on this page is the real trained model from both papers. It runs inside your browser, and its two convolution layers run in WebAssembly. Every number in the paper sections comes from the papers themselves.
The 30-second version
- The problem. Highlights can look right without showing what the model actually used. Looking right is called plausible. Showing what the model used is called faithful.
- The test. Delete what the highlighter coloured. If the model's answer collapses, the highlight was faithful. If the model shrugs, it wasn't.
- Paper B (first). One model (MalConv), twelve highlighters. The built-in "XAI point" is cheap and better than random, but LIME, LEMNA-like and Partition SHAP are clearly more faithful. An LLM that only sees the verdict is the least faithful.
- Paper A (latest). Same core, plus two more models (a character n-gram model and CodeBERTa), stronger tests, 82 authors, three LLMs, and a new question: what would a human analyst learn from the highlights? Answer: the right thing about the model, but not about the authors.
1 · The question: plausible or faithful?
Dr. Nassar and Bou Abdo built a tool called "Highlight to explain". A neural network called MalConv reads code as raw bytes and classifies it. They added a cheap built-in highlighter, the XAI point, and showed it in an editor plugin. To check it, they used a Python 2 vs Python 3 task and graded the highlights against patterns a human expects, like print "hello" (Python 2 style).
That grading checks whether the highlight agrees with a human. It doesn't check whether the model used those characters. A model can be right for reasons nobody expected, and some highlighters return roughly the same picture no matter what model you attach them to.
The thesis moves to a task where no human can supply the answer key: which of 20 programmers wrote this Python file? No one can mark the characters that make a file "belong" to its author. So the only thing you can check is faithfulness, by poking the model.
2 · The data: 20 Code Jam programmers
Google Code Jam was a yearly programming contest. A public archive holds 378,487 Python solutions to 96 problems from 2017 to 2020. Both papers keep the 20 authors who solved the most distinct problems (with at least 20 training, 4 validation and 8 test files each) and remove exact duplicates.
Why the split is by problem, not by file
If the same problem appears in training and testing, a model can cheat: it recognises the problem ("this is the pancake-flipping one, and Author 3 solved that") instead of the author's style. So whole problems are assigned to one side. Every test file is a problem the model has never seen. A control that splits the same amount of data randomly by file reaches 91.6%, so the problem split isn't hiding a big leak.
3 · MalConv, layer by layer (running live)
MalConv was built to spot malware in raw program bytes. Nassar modified it and both papers resize it for 20 authors. It has 28,244 learned numbers (parameters) and reaches 96.1% test accuracy on the seed used below (94.4% ± 1.6 over three seeds). Here is the whole path from a file to a verdict.
Bytes: the file becomes numbers 4096 integers, 0–255
A computer stores text as bytes. Each character becomes a number from 0 to 255: p is 112, a space is 32, a newline is 10. MalConv reads the first 4,096 bytes of the file. Shorter files are padded with zeros, longer ones are cut. There is no notion of words, keywords or syntax here. The model sees one long row of numbers.
Embedding: each byte becomes 8 numbers 4096 × 8
A byte value like 112 is just an ID. "112 is bigger than 32" means nothing. So the model keeps a lookup table with 256 rows (one per possible byte) and 8 columns. Each byte is replaced by its row: 8 numbers that describe it. That's 256 × 8 = 2,048 learned numbers.
Nobody chose these numbers. Training adjusted them until bytes that play similar roles in an author's style ended up with similar rows. You can check that below: pick a character and see which other characters the model learned to treat as its neighbours.
The whole file as an 8-row picture
Each column is one byte of the file (first 512 shown), each row one of the 8 numbers. Blue is positive, red is negative. Repeated patterns in the code show up as repeated stripes.
Convolution: 128 pattern detectors slide along the file two maps, each 4096 × 128
A filter is a small stencil 8 bytes wide. At every position it looks at the 8 bytes around it (3 before, the byte itself, 4 after), multiplies each of their 8 embedding numbers by one of its own 64 weights, adds everything up and adds a bias. A big result means "the pattern I'm tuned for is here". The filter then moves one byte to the right and does it again, 4,096 times.
There are 128 filters, and MalConv has two such layers side by side reading the same input: one says what it found, the other decides how much of it to let through (step 4). That's 2 × 4096 × 128 = over a million results, each a sum of 64 products. This is the heavy part, and it's the part running in WebAssembly in your browser.
The gate: a dimmer switch on every detector 4096 × 128
The second convolution goes through a sigmoid, a squashing function that turns any number into something between 0 and 1. That value multiplies the first convolution's result. So at each position, each filter's finding is let through fully (×1), dimmed, or switched off (×0) depending on context. This is called a gated convolution.
For the filter chosen in step 3, across the whole file. Hover to read values.
ReLU + L1: the "XAI point" 4096 × 128, mostly zeros
ReLU sets every negative number to 0 and keeps positive ones. Nassar's idea was to put an L1 penalty right here during training: the model pays a small cost (λ = 0.0001) for every unit of activation. To keep the cost low, the model learns to fire only where it really matters. That turns this map into a built-in highlighter, the XAI point: the positions that light up are "what the model noticed".
With the penalty, only about 2% of these values are non-zero (against 49% without it). Your file right now: … non-zero.
Rows are the 128 filters, columns are positions in the file. Darker means stronger. Nearly everything is blank, which is the point.
The XAI point as a highlight
Each position's strongest filter (max over the 128) is spread over the 8 bytes it read, then each code token takes its strongest byte. Darker orange = higher score.
Global max-pool: "did this pattern appear anywhere?" 128 numbers
For each filter, the model keeps only its single highest value over the whole file, and forgets where it was. 4096 × 128 numbers become just 128. This makes the model blind to position: a habit at the top of the file counts the same as at the bottom.
This is also why the XAI point can be misleading: one filter's peak decides everything for that filter, and the other 4,095 positions are thrown away.
The filters that fired hardest on this file, and what they fired on
Dense layer: a weighted vote 128 → 64
Each of 64 "hidden" units takes all 128 pooled numbers, multiplies each by its own weight and adds a bias: 64 × 128 + 64 = 8,256 learned numbers. Think of each unit as a judge who weighs the 128 pieces of evidence differently.
As in Nassar's original, there is no activation between this layer and the next. Two linear steps in a row collapse into one, so mathematically the head is a single 20 × 128 table: "how much each filter votes for each author". Paper B checked that adding a ReLU here changes nothing (same 96.1% and the same deletion scores). This collapse is what makes the class-aware XAI point possible (section 4).
Output + softmax: one score per author, then probabilities 64 → 20
20 more weighted sums (64 × 20 + 20 = 1,300 numbers) give one raw score, a logit, per author. Softmax turns them into probabilities that add up to 100%: raise e to each score and divide by the total. The biggest wins.
How the 28,244 numbers were learned (training)
Start with random numbers. Show the model a batch of 16 training files. For each, compare its probabilities with the true author using cross-entropy (a big penalty when it gives the true author a low probability), add the L1 penalty from step 5, and nudge every number slightly in the direction that lowers the total. The nudging rule is Adam with learning rate 0.001. Repeat over all training files for up to 40 passes (epochs), and stop early if validation accuracy hasn't improved for 8 passes. λ was picked from {0, 10⁻⁵, 10⁻⁴, 10⁻³, 10⁻²} by validation accuracy among penalised models.
| Layer | Shape | Learned numbers |
|---|---|---|
| Embedding | 256 × 8 | 2,048 |
| Convolution 1 (what) | 128 filters × 8 bytes × 8 + 128 | 8,320 |
| Convolution 2 (gate) | 128 filters × 8 bytes × 8 + 128 | 8,320 |
| Dense | 128 → 64 | 8,256 |
| Output | 64 → 20 | 1,300 |
| Total | 28,244 |
4 · How a highlight is made: the explainers
An explainer (highlighter) takes the model and a file and gives every code token a score. A token is a word-like run (input, range, T) or a single symbol ((, =). All explainers score the same tokens, so they can be compared fairly. To "remove" a token, its characters are replaced by spaces, so every other byte stays in place and the convolution windows don't shift.
Four you can run right here
Darker orange = higher score. The top 10% of tokens (what the editor plugin would highlight) are underlined.
XAI point (Nassar's)
Score = the strongest filter activation at that spot (step 5). It's free: one pass through the model. But it is class-agnostic: it never looks at which author is being explained. It shows "what lit up", not "what pointed to Author 7".
XAI point, class-aware (thesis)
Because the head is one linear table (step 7), the score for author c is exactly the sum of each filter's peak times that filter's vote for c. So credit each filter's peak × vote back to the spot where it peaked. Same cost, one pass, but now it knows which author it's explaining, and evidence against the author gets a negative score.
Token occlusion
Hide one token at a time and watch the author's probability. The drop is the token's score. Honest and simple, but needs one model run per token.
Random
Random scores. Any explainer worth using must beat this by a lot. It's the floor.
The others in the papers (in plain words)
| Explainer | How it works | Cost |
|---|---|---|
| Grad-CAM | Uses the gradient (how much the author's score would change) at the XAI-point layer to weight each filter's map, keeping only positive parts. | one forward + one backward pass |
| Gradient × input | The gradient with respect to each byte's 8 embedding numbers, times the numbers themselves. | one backward pass |
| Integrated gradients | Slowly morph the file from all-spaces into the real file in 32 steps, add up the gradients along the way. Fixes gradient × input's blind spots. | 32 passes, ~0.4 s |
| Sliding window | Nassar's other method: blank an 8-byte window, slide it 4 bytes, measure the drop. | ~1,000 passes |
| LIME | Make 1,000 copies of the file with random tokens hidden, ask the model about each, then fit a simple straight-line model that predicts the answer from "which tokens were kept". Its weights are the scores. | 1,000 model calls, ~6 s |
| LEMNA-like | LIME plus a "fused lasso" rule that pushes neighbouring tokens to get similar scores, so it highlights whole stretches. Same 1,000 samples. | 1,000 calls, ~10 s |
| Kernel SHAP | Game theory (Shapley values): a token's score is its average contribution over many random coalitions of other tokens. Estimated with 1,000 calls. | 1,000 calls |
| Partition SHAP | Shapley values over a hierarchy (groups of neighbouring tokens split recursively), which spends the 1,000 calls far more efficiently on text. | 1,000 calls, ~5 s |
| LLM | Claude Sonnet 5 (Paper A adds GPT-5.6 and Gemini Pro) gets the file with line numbers, the predicted author and three other files by that author, and ranks up to 15 lines that "reveal the author". It never sees the classifier, like an assistant asked to justify another tool's verdict. | one API call |
| Model-free reference | Ranks tokens by how typical they are of the author in the training data. Ignores the model completely. Shows what "plausible" looks like. | free |
| Greedy deletion | (Paper B only) Directly searches for the deletion order that kills the prediction fastest. A reference, not a real explainer. | ~20 s |
5 · How a highlight is tested
The main test: delete and watch (ABPC)
Sort the tokens by score. Hide the top 1%, 2%, 5% … 100% and record the probability of the predicted author each time. That's the MoRF curve (most relevant first). It should crash fast. Then do it in reverse, least relevant first: the LeRF curve. It should stay high for a long time. ABPC is the area between the two curves. Big gap = faithful. A random ranking gives about 0.
This is one file, so it's a demo of the method, not a result. The papers average over all 254 test files.
Three quick numbers at 10% (what the plugin would show)
- Comprehensiveness: how much does the probability drop if you hide the top 10%? (high is good)
- Sufficiency: how much does it drop if you keep only the top 10%? (low is good)
- Flip rate: how often does hiding the top 10% change the predicted author? (high is good)
Randomization test: does the highlight even depend on the model?
Scramble the model's weights from the output layer downwards and compare the highlights before and after. A faithful explainer's map should change. Try it on the head:
The class-agnostic XAI point is computed before the head, so it can't change at all (correlation 1.00). That's exactly what Paper B found and why the class-aware version was added.
Retraining check (ROAR)
Deleting tokens creates weird files the model never saw in training, which might fool the deletion test. So: remove each explainer's top tokens from every file, retrain a new model from scratch, and see how much accuracy falls. If the explainer found real evidence, the new model has less to learn from. Explanations are "cross-fitted": no model explains its own training files.
Planted identifiers: the one answer key we can build
Give each author a made-up 8-letter name (say qvxmetrz) and insert a line qvxmetrz = 0 into their files. Now we know a piece of evidence exists. But does the model use it? That's measured directly, never from the explainers: blank the identifier or swap in another author's, and see if the prediction drops. Then a faithful explainer should rank the identifier high exactly in the files where the model relies on it, and not elsewhere.
Statistics, simply
Files by the same author aren't independent, so intervals come from resampling authors, not files (a cluster bootstrap). Pairs of explainers are compared with an exact sign-flip test (one sign per author), corrected for running many comparisons (Holm). Explainers sharing a letter in the results table are not significantly different.
6 · First, reproducing Nassar's paper
Before changing anything, both papers rebuild the original setup: the small MalConv (window 7, 64 filters, λ = 0.01, 8 dense units, one yes/no output) on Nassar's released Python 2 vs 3 dataset of single lines, and also load the authors' own released model.
- It reproduces. The released model gets 99.3% on its training lines; retrained models get 96.7% on lines they never saw. The XAI point "localises" the Python 2 pattern in 90.6% of lines under the released grading rule, inside the published 80–98% range.
- A catch in the grading. The rule accepts any highlight containing a
ustring prefix, and 85% of the annotations are that one pattern. On the rarer patterns, with a strict rule, the released model localises 88.7% but retrained models only 58.3%, with big swings between seeds. - A first faithfulness test. Blanking the XAI point's highlighted window flips the prediction in 64.0% of held-out lines, against 9.9% for a random window of the same size. Much better than chance, but in about a third of lines the model doesn't need what was highlighted.
Paper B
"Highlight, then verify". 8 pages. One model (MalConv), twelve explainers, one LLM. About 10 rounds of review by several Claude models under Dr. Safa's zero-bad-feedback rule, ending at weak accept. Frozen at git tag planB-weak-accept.
The 28 later "what-if" versions (v4–v28) used made-up numbers to see what reviewers would want. They were never real results.
Paper A
9 pages, built 2–3 Oct 2026 on branch yield. You approved making the strongest what-if claims real, except C++, a GitHub dataset, Qwen and a 30-file selector. Every number comes from a real run, mostly on pc00. Not yet through the review loop.
Rule of the branch: no hypothetical numbers, and LLM steps use stored Claude outputs or non-Claude models.
MalConv only gets the full battery. A character n-gram model (96.9%), a random forest on style features (94.9%) and CodeBERTa (92.5%) appear only as accuracy references.
Three model families get the full battery: MalConv, the character n-gram model and CodeBERTa. The random forest stays an accuracy reference. Their mechanisms:
How the character n-gram model works
1. Cut the file (first 4,096 bytes) into every overlapping chunk of 1, 2, 3 and 4 characters. print gives p r i n t pr ri in nt pri rin int prin rint.
2. Keep the 200,000 most common chunks seen in at least 2 training files. Each file becomes a list of 200,000 numbers, one per chunk, using TF-IDF: how often the chunk appears in this file (log-scaled), times how rare it is across files (common chunks like a space count for little).
3. Logistic regression: for each of the 20 authors, one weight per chunk. Score = sum of chunk value × weight, then softmax. The regularisation strength C is picked from {1, 10, 100} on validation. No layers, no training randomness: same data, same model.
Its weights are a built-in explainer ("LR coefficients"): a token's score comes from the weights of the chunks it contains.
How CodeBERTa works
CodeBERTa-small is a transformer (the same building block as ChatGPT, but small and made for reading code, not writing it), pre-trained by Hugging Face on about 2 million functions from GitHub (CodeSearchNet), then fine-tuned here on the 20 authors.
- Tokenizer (BPE): code is split into pieces from a vocabulary of 52,000 common chunks, like
def,range,_input. It reads the first 512 pieces of the first 4,096 bytes (35% of test files are cut there). - Embedding: each piece becomes 768 numbers, plus 768 more that encode its position.
- 6 transformer layers, each with:
- Self-attention, 12 heads: every piece looks at every other piece and decides which ones matter to it ("this
)belongs to that("), then mixes in their information. 12 heads = 12 different ways of looking. - Feed-forward: each piece's 768 numbers go through a 3,072-wide layer and back.
- Shortcuts (residuals) and normalisation keep the numbers stable.
- Self-attention, 12 heads: every piece looks at every other piece and decides which ones matter to it ("this
- Classification head: the first piece's final 768 numbers (a summary slot) go through a small dense layer to 20 scores, then softmax.
Fine-tuned with AdamW, learning rate 5×10⁻⁵. Unlike MalConv, it knows order and context everywhere, not just inside 8-byte windows.
ABPC on the 254 test files, with 95% intervals (hover for cost and group). The leaders are LEMNA-like, LIME and Partition SHAP. Nassar's XAI point is far above random but significantly below them.
Same table, same numbers (Paper A reuses Paper B's MalConv results). Then it adds a check Paper B didn't have:
ROAD: deletion without weird files
Instead of replacing a removed token with spaces, a language model fills in a different plausible token. The files stay realistic. ABPC falls for everyone but the order barely moves (rank correlation 0.97 with the space version):
So the "weird files" worry about space masking doesn't drive the ranking.
Faithfulness vs seconds per file (log scale, 4 CPU cores). Integrated gradients gets most of the leaders' faithfulness in 0.4 s against 5–10 s. Grad-CAM and the class-aware XAI point cost one pass. Recommendation: integrated gradients or Grad-CAM in an editor; LIME-type explainers for offline audits.
Same cost picture. Paper A's conclusion names integrated gradients as the explainer it would show in the editor for MalConv.
Other seeds and a second pool of 20 authors (criteria fixed in advance): LEMNA-like 0.768, LIME 0.762 and Partition SHAP 0.750 lead again; the XAI point (0.647) trails again; the LLM (0.500) is last again.
Same seeds and second pool, plus 82 authors (870 test files, after removing cross-author duplicates, reused files, files with the author's handle and files under 200 bytes). Accuracy: MalConv 89.9% (5 seeds), n-gram 95.1%, CodeBERTa 90.8%. More authors don't help the deep models. Ranking correlation with 20 authors: 0.92.
And on the other models, the leader changes:
CodeBERTa
n-gram model
Partition SHAP leads on CodeBERTa, token occlusion (and the model's own weights) on the n-gram model. An explainer validated on one model can't be assumed faithful on another. (The n-gram model's scores are lower overall because it spreads its evidence over 200,000 features; one token rarely matters much.)
Randomization: the class-agnostic XAI point doesn't change at all when the head is scrambled (correlation 1.00): it can't depend on which author is explained. Grad-CAM keeps 0.51 because it keeps only positive scores. The test can't separate the rest well on this architecture.
Retraining (ROAR), accuracy after removing each explainer's top tokens and retraining (no removal: 94.3%). * = significantly below random.
Only at 40% removal do the good explainers separate from random. At 10–20% the retrained model copes.
Same MalConv results, now also on the other families:
- Randomization: every CodeBERTa explainer drops to |ρ| ≤ 0.09; n-gram explainers to |ρ| ≤ 0.03 once the weights are random. They all depend on the model.
- Retraining, n-gram: removing 40% of the tokens chosen by its own weights, LEMNA-like, LIME or occlusion leaves 66.5–71.7% accuracy, against 96.9% for random removal. Even at 10%: 91.7–93.3% vs 96.9%.
- Retraining, CodeBERTa: recovers more easily. At 40%, occlusion's tokens leave 87.2%, integrated gradients' 90.9%, random 93.7% (two seeds).
120 files, identifier in every file. The model relies on it heavily in only 16. Model-based explainers' ranking of the identifier follows that reliance (ρ 0.82–0.92, Kernel SHAP 0.70). The LLM ranks it first more often (68% of files) because it's the most author-typical string, but doesn't follow the model at all (ρ = 0.09).
When the model barely uses the identifier, the LLM still picks its line in 46% of files if the reference files contain it, and 0% if they don't. It finds it by matching the references, not the classifier.
Rerun on all three families, 254 test files each, with a permutation null. On MalConv, 70 files rely heavily on the identifier.
On MalConv every model-based explainer tracks reliance (ρ 0.59–0.75; null's 95th percentile 0.11) and the model-free baseline doesn't (0.39). On the n-gram model and CodeBERTa the identifier is so author-typical that the model-free baseline tracks reliance too. So this control catches a blind explainer but can't rank good ones there.
Claude Sonnet 5 only. Most readable, least faithful (ABPC 0.492). Not just a format problem: LIME forced to rank 15 whole lines still beats it. Asking "which lines did the classifier rely on" changes nothing (0.496); Claude Opus 5.5 is barely better (0.517).
Given the classifier's outputs (the probability when each line is blanked), it jumps to 0.591 and its flip rate from 22% to 51%, but still trails simply sorting those numbers.
Three LLMs, all from different companies, all 254 files:
Repeat runs vary little (0.518 vs 0.512), and a neutral prompt changes little (0.489). What the LLMs lack is access to the model. With it they reach 0.59–0.60, but never beat just sorting the numbers they were handed (0.625). Their only extra is a readable sentence, which wasn't tested with users.
Does the L1 penalty help? Ten seeds per λ. The penalty slightly raises the XAI point's normalised ABPC (0.707 at λ=0 to 0.779 at 10⁻³) while accuracy falls (96.6% to 91.0%). It sparsifies the map but doesn't close the gap to the leaders.
What the explainers highlight: about a fifth of top tokens on input/output handling (stdin, readline, split, format) and a similar share on each author's scaffolding (imports, main guard, test-case loop), against 6% and 8% for random. Nassar's XAI point spends over half its top tokens on punctuation, about as much as random.
What would an analyst learn from the highlights?
Token-by-token scores say "these characters matter". An analyst wants a sentence: "the model recognises this author by how they read input". Paper A builds those sentences mechanically and tests them.
- Label every line with one of 16 statement types using Python's own tokenizer: header, comment, import, input reading, output writing, the test-case loop (the loop that prints
Case #), function/class definitions, other loops, branches, assignments, augmented assignments, returns, other control flow, calls, other. - Turn highlights into hypotheses. For each correctly attributed file, take the explainer's top 10% tokens. For each type: share of top tokens on that type's lines, minus that type's share of all tokens ("excess share"). Average over files and rank the types. The explainer's hypothesis = its top-3 types.
- Test the hypothesis: remove every line of those types from the test files and measure the accuracy drop.
- Fair comparison: removing more code always hurts, so each removal is paired with a matched random control that removes the same number of random lines per file. Excess drop = drop(types) − drop(matched random).
- Measure it twice: on the original model (what does this model use?) and after retraining on the edited files, 5 seeds (does the task need it?).
The removal rule, the matched control and the pass rule were written down before any explainer's result was computed.
The hypotheses
On MalConv, all eleven model-based explainers say: output writing first, the test-case loop second. Third differs (input reading for Partition SHAP, Grad-CAM and the XAI point; imports for LIME and LEMNA-like). CodeBERTa's explainers say imports, function definitions and input reading or headers. The n-gram's say output writing, input reading, imports.
Result 1: on the original model, the hypotheses are right
Removing Partition SHAP's top 3 types drops MalConv from 96.1% to 64.6%, against 85.0% for the matched control: 20.5 points excess. The other model-based explainers give 13.4–16.1. The random explainer's hypotheses fall 6.3 short of their control, the model-free baseline 15.0 short, and a mutual-information baseline (what's informative in the data) is only 0.8 above. Only model-based explainers find what this model uses.
Result 2: the hypotheses are model-specific
Excess drop (points) when the top-3 types from one model's explainers are removed from another model's test files.
CodeBERTa's own hypotheses: 19.3; MalConv's applied to CodeBERTa: 0.4. The n-gram model barely reacts to any removal: with 200,000 features it doesn't depend on any one statement type.
Result 3: retraining erases it
Retrained on the edited files, every model-based explainer's top-3 types cost no more than random lines (excess −1.8 to −0.1). Even removing all five template types together leaves a retrained MalConv at 84.6% (control 86.1%), while the original model falls to 53.5%. Only explainers that propose input reading pass the pre-registered rule, by 6.3 points [2.9, 9.8].
The change of plan (deviation 1), and why it's a finding
The original plan was to rewrite a statement type without changing behaviour (canonical layout + consistent renaming, verified by token-stream equality; 7 of 889 files couldn't be verified and stayed unchanged). Rewriting every line drops the original MalConv from 96.1% to 77.6%, but a model retrained on rewritten files gets 92.8% vs 94.3%. The n-gram model loses 1.2 points, CodeBERTa 1.0. No single type costs more than 1.2.
So the operator couldn't separate any hypothesis from noise, and it was swapped for removal before any explainer's yield was computed, logged as deviation 1 in PLAN.md. The rewrite result stays in the paper as a finding: these models use layout and naming, but the task doesn't need them.
Validate the explainer on the model it explains (for instance a quick deletion test against random) before showing highlights as grounds for trust. XAI point: cheap, better than random, less faithful. Class-aware XAI point: same cost, narrows the gap. LLM without model access: plausible, not faithful.
All of Paper B's verdict, holding under ROAD, on more seeds and on 82 authors, plus: the best explainer depends on the model; highlights should be presented as evidence about the model, not about the authors; and giving an LLM the model's numbers helps but never beats the numbers themselves.
8 · Side by side
| Paper B (first) | Paper A (latest) | |
|---|---|---|
| Question | Are the highlights faithful? Which explainer, at what cost? | Same, plus: does it depend on the model, and what would an analyst learn? |
| Models with full tests | MalConv | MalConv, n-gram LR, CodeBERTa |
| Authors | 20 (+ a second pool of 20) | 20 (+ second pool) + 82 |
| Explainers | 12 incl. 1 LLM | 12 incl. 3 LLMs (Claude, GPT-5.6, Gemini) |
| Deletion tests | ABPC with spaces | ABPC with spaces + ROAD (infill) |
| Retraining / randomization / planted | MalConv | All three families |
| Statement-type test | not in this paper | Removal + matched control + cross-model matrix + retraining |
| L1 penalty study | Yes, 10 seeds per λ | dropped for space |
| Leader on MalConv | LEMNA-like, LIME, Partition SHAP | Same |
| Leader elsewhere | not tested | Partition SHAP on CodeBERTa, occlusion on n-gram |
| Status | ~10 review rounds, weak accept, frozen | Real numbers, not yet reviewed |