Abstract
The rapid proliferation of generative models raises the model attribution problem: given only an image, can we determine which model produced it? We propose a lightweight CNN to solve this problem in a strict black-box setting. The CNN operates on multiple image patches to handle varying image size and improve accuracy.
It attributes more models at higher accuracy than prior work, reaching 98.9% on 25-class DRAGON and 95.0% on 27-class OpenFake; runs in a few milliseconds at a cost nearly independent of candidate-set size; and remains robust to transformations encountered in the wild.
Beyond closed-set attribution, its learned representation serves as a reusable fingerprint backbone: it supports open-set rejection and few-shot enrollment, recovers model-family structure without lineage supervision, and groups images from unseen generators by source. Controlled ablations and causal perturbations show that its attribution decisions are driven by a spatially local, low-level signal that behaves as a generator-specific fingerprint.
Method
Fixed-size RGB patches, overlapping at the edges so every pixel is covered. The model is resolution-invariant, from 64² to 4096² without retraining.
A compact 5.9M-parameter CNN scores each patch against every candidate generator in one forward pass — image-only, black-box, no model internals required.
Per-patch logits combine into one image-level label via an overlap-corrected weighted average.
Closed-set attribution
Evaluated on DRAGON (25 generators) and OpenFake (27 generators), from the image alone — no model access at all.
| Method | Params (M) | Infer. (ms) | DRAGON (25) | OpenFake (27) |
|---|---|---|---|---|
| DE-FAKE | 151 | 9.96 ± 1.31 | 62.0 | 57.3 |
| USIA | 427 | 9.61 ± 1.01 | 52.1 | 50.1 |
| LIDA | 23.5 | 23.6 ± 12.5 | 24.0 | 17.7 |
| OCC-CLIP | 151 | 7.87 ± 1.12 | 8.6 | 15.8 |
| Ours | 5.9 | 2.79 ± 0.04 | 98.9 | 95.0 |
On AEDR's eight-model benchmark, where competitors get each candidate's autoencoder, our classifier reaches higher mean pairwise accuracy at two orders of magnitude lower inference cost, image-only.
| Method | Access | Infer. (s) | Acc. (%) |
|---|---|---|---|
| LatentTracer | Model weights | 24.06 | 70.3 |
| AEDR | VAE weights | 0.267 | 95.1 |
| Ours | Image only | 0.0028 | 99.5 |
Open-set & adaptation
Trained on 17 of OpenFake's 27 generators, the classifier can flag sources it has never seen, and quickly incorporate new ones once labels arrive.
Images from the 10 held-out sources are flagged as unknown by thresholding the classifier's own confidence — no calibration set, no auxiliary outlier data.
The backbone already learns a general fingerprint space, so a generator it has never seen is admitted by freezing that backbone and fitting only a linear head — no retraining. Even a single labeled image per class beats the published baselines, and the head fit takes under four seconds.
| Method | 1-shot | 10-shot |
|---|---|---|
| ResNet | 17.4 | 21.4 |
| DIRE | 14.3 | 17.2 |
| ESSP | 17.0 | 22.4 |
| LIDA | 40.4 | 54.0 |
| Ours (OpenFake backbone) | 47.7 ± 3.5 | 72.0 ± 0.8 |
| Ours (DRAGON backbone) | 52.0 ± 3.6 | 72.4 ± 0.9 |
Universal feature extractor
Trained only to classify known generators, the network's features generalize beyond that task: they organize sources it has never seen, and recover the lineage of the ones it has, with no fine-tuning, labels, or calibration.
UMAP of penultimate features for ten generators never seen in training. With no target count, density-based clustering auto-estimates seven clusters against the true 10 sources (ARI 0.64, NMI 0.84); the only collapses are within a shared model family.
BibTeX
Preprint on arXiv; conference citation will be updated on acceptance.