← Paper page CREDIT: Certified Ownership Verification of Deep Neural Networks Against Model Extraction Attacks Ask this paper arXiv ↗

CREDIT: Certified Ownership Verification
of Deep Neural Networks
Against Model Extraction Attacks


If someone copies your model just by querying it, can you prove the copy came from you, with guaranteed error rates?

Bolin Shen, Zhan Cheng, Neil Zhenqiang Gong, Fan Yao, Yushun Dong · ICML 2026

Press → or click to step through

Stealing a model through its front door

customers ask, get answers your model served through an API attacker many queries input–output pairs trains a working copy

The copy closely mimics the original, at a small fraction of the cost of building it from scratch.

This is a model extraction attack. The original took careful design, proprietary data, heavy computing and long training.

Is this suspect model a copy of mine?

your model related? ? a suspect model independent built from scratch a copy extracted from your API mistake: false alarm mistake: missed copy

Existing defences plant watermarks, which can cost accuracy, or compare fingerprints, which need many extra trained models. Neither comes with rigorous guarantees.

CREDIT’s goal: a yes-or-no answer with a proven limit on each kind of mistake.

Step 1: blur the model’s inner signals a little

input your model inner signal (embedding) + small random noise = protected signal used to answer queries ceiling no other model can share more than this similarity with your model

The noise puts a provable ceiling on how much any other model can share with yours (Thm 3.1), and the paper shows this ceiling is essentially tight (Thm 3.2).

Similarity is measured by mutual information: how much one model’s inner signals reveal about the other’s. The noise is the Gaussian mechanism from differential privacy. More noise means a lower ceiling but some lost accuracy; the paper picks the noise level by balancing the two (Section 3.3).

Step 2: score the suspect and compare with a threshold

no claim declared a copy ceiling 0 similarity score → independent models score near zero copies trained to imitate, so they score high CREDIT threshold set higher if the attacker may make more queries; lower with a bigger check set or more signal dimensions

Certified: the chance of a false alarm and the chance of a missed copy each have a proven upper limit, and both shrink exponentially as the check set grows (Thm 3.4).

The guarantee holds under the paper’s conditions: an independent model’s expected score sits below the threshold (ideally near zero), and a copy trained with enough queries scores above it. Scores are estimated from nearest neighbours (the KSG estimator) on a held-out check set. Positions on this scale are schematic.

What does CREDIT decide?

no claim declared a copy threshold 0.59 ceiling 4.17

Values come from different experiments in the paper, placed on one scale: threshold 0.59 and ceiling 0.59 + 3.58 at the chosen noise level (Fig. 1c); independent model (App. D.1, four estimator settings); copies (Table 8).

Copies were told apart perfectly in every reported test

how well each method tells copies from independent models (AUC, CIFAR-10) 50 = coin flip 100 = perfect
network

CREDIT scored a perfect 100 on all four networks; the watermark and fingerprint baselines ranged from 40.74 to 80.25.

Table 2. Verification also stayed perfect against worst-case disguise attacks (Table 8) and on a Word2Vec text model (Table 9).

Little accuracy cost, much faster checks

accuracy of the protected model (%)
seconds to check one suspect

No extra models to train: CREDIT just estimates one similarity score and compares it with the threshold.

Left: CIFAR-10 with ResNet (Table 1). Right: ResNet-50 (Table 3); on VGG-16, 22.23 s versus 72.74–83.10 s. Preparing CREDIT takes 0.001 s. On graph data, CREDIT was the most accurate defence in every case tested.

What this means for model owners

  1. 1Ownership checks can come with guaranteesCREDIT bounds the chance of a false alarm and of a missed copy, under the paper’s stated conditions; both bounds shrink as the check set grows.
  2. 2A little noise makes similarity measurableNoise on the model’s inner signals caps how much any model can share with it, while copies still score well above the threshold.
  3. 3Strong in practicePerfect separation in the paper’s CIFAR-10 tests, accuracy close to the undefended model, and checks in about 22 seconds.

Paper page · arXiv · Code