CREDIT: Certified Ownership Verification
of Deep Neural Networks
Against Model Extraction Attacks
If someone copies your model just by querying it, can you prove the copy came from you, with guaranteed error rates?
Bolin Shen, Zhan Cheng, Neil Zhenqiang Gong, Fan Yao, Yushun Dong · ICML 2026
Press → or click to step through
Stealing a model through its front door
The copy closely mimics the original, at a small fraction of the cost of building it from scratch.
This is a model extraction attack. The original took careful design, proprietary data, heavy computing and long training.
Is this suspect model a copy of mine?
Existing defences plant watermarks, which can cost accuracy, or compare fingerprints, which need many extra trained models. Neither comes with rigorous guarantees.
CREDIT’s goal: a yes-or-no answer with a proven limit on each kind of mistake.
Step 1: blur the model’s inner signals a little
The noise puts a provable ceiling on how much any other model can share with yours (Thm 3.1), and the paper shows this ceiling is essentially tight (Thm 3.2).
Similarity is measured by mutual information: how much one model’s inner signals reveal about the other’s. The noise is the Gaussian mechanism from differential privacy. More noise means a lower ceiling but some lost accuracy; the paper picks the noise level by balancing the two (Section 3.3).
Step 2: score the suspect and compare with a threshold
Certified: the chance of a false alarm and the chance of a missed copy each have a proven upper limit, and both shrink exponentially as the check set grows (Thm 3.4).
The guarantee holds under the paper’s conditions: an independent model’s expected score sits below the threshold (ideally near zero), and a copy trained with enough queries scores above it. Scores are estimated from nearest neighbours (the KSG estimator) on a held-out check set. Positions on this scale are schematic.
What does CREDIT decide?
Values come from different experiments in the paper, placed on one scale: threshold 0.59 and ceiling 0.59 + 3.58 at the chosen noise level (Fig. 1c); independent model (App. D.1, four estimator settings); copies (Table 8).
Copies were told apart perfectly in every reported test
CREDIT scored a perfect 100 on all four networks; the watermark and fingerprint baselines ranged from 40.74 to 80.25.
Table 2. Verification also stayed perfect against worst-case disguise attacks (Table 8) and on a Word2Vec text model (Table 9).
Little accuracy cost, much faster checks
No extra models to train: CREDIT just estimates one similarity score and compares it with the threshold.
Left: CIFAR-10 with ResNet (Table 1). Right: ResNet-50 (Table 3); on VGG-16, 22.23 s versus 72.74–83.10 s. Preparing CREDIT takes 0.001 s. On graph data, CREDIT was the most accurate defence in every case tested.
What this means for model owners
- 1Ownership checks can come with guaranteesCREDIT bounds the chance of a false alarm and of a missed copy, under the paper’s stated conditions; both bounds shrink as the check set grows.
- 2A little noise makes similarity measurableNoise on the model’s inner signals caps how much any model can share with it, while copies still score well above the threshold.
- 3Strong in practicePerfect separation in the paper’s CIFAR-10 tests, accuracy close to the undefended model, and checks in about 22 seconds.
Paper page · arXiv · Code