Scalable Black-Box Model Attribution for Images

Abstract

The rapid proliferation of generative models raises the model attribution problem: given only an image, can we determine which model produced it? We propose a lightweight CNN to solve this problem in a strict black-box setting. The CNN operates on multiple image patches to handle varying image size and improve accuracy.

It attributes more models at higher accuracy than prior work, reaching 98.9% on 25-class DRAGON and 95.0% on 27-class OpenFake; runs in a few milliseconds at a cost nearly independent of candidate-set size; and remains robust to transformations encountered in the wild.

Beyond closed-set attribution, its learned representation serves as a reusable fingerprint backbone: it supports open-set rejection and few-shot enrollment, recovers model-family structure without lineage supervision, and groups images from unseen generators by source. Controlled ablations and causal perturbations show that its attribution decisions are driven by a spatially local, low-level signal that behaves as a generator-specific fingerprint.

Same prompt, twenty-five generators. A single prompt rendered by twenty-five text-to-image models.

Method

Three steps, one small network

01 / Patch division

Split into 64×64 patches

Fixed-size RGB patches, overlapping at the edges so every pixel is covered. The model is resolution-invariant, from 64² to 4096² without retraining.

02 / Per-patch classification

Classify each patch from raw pixels

A compact 5.9M-parameter CNN scores each patch against every candidate generator in one forward pass — image-only, black-box, no model internals required.

03 / Aggregation

Combine per-patch votes

Per-patch logits combine into one image-level label via an overlap-corrected weighted average.

Closed-set attribution

Black-box and white-box accuracy

Black-box comparison

Evaluated on DRAGON (25 generators) and OpenFake (27 generators), from the image alone — no model access at all.

98.9%
Top-1 accuracy · 25-class DRAGON
95.0%
Top-1 accuracy · 27-class OpenFake
2.8ms
Inference time per image
Table 1 — Black-box setting. Top-1 accuracy (%) on 25-class DRAGON and 27-class OpenFake. From the image alone, a compact CNN attains the highest accuracy at the largest class count, at the lowest inference cost. Inference times re-measured on our hardware.
MethodParams (M)Infer. (ms)DRAGON (25)OpenFake (27)
DE-FAKE1519.96 ± 1.3162.057.3
USIA4279.61 ± 1.0152.150.1
LIDA23.523.6 ± 12.524.017.7
OCC-CLIP1517.87 ± 1.128.615.8
Ours5.92.79 ± 0.0498.995.0

White-box comparison

On AEDR's eight-model benchmark, where competitors get each candidate's autoencoder, our classifier reaches higher mean pairwise accuracy at two orders of magnitude lower inference cost, image-only.

Table 2 — White-box setting. Mean pairwise accuracy and per-image inference time on AEDR's eight-model benchmark. Our method outperforms the alternatives while being two orders of magnitude faster and strictly black-box. Baseline accuracies from AEDR; all inference times re-measured on our hardware.
MethodAccessInfer. (s)Acc. (%)
LatentTracerModel weights24.0670.3
AEDRVAE weights0.26795.1
OursImage only0.002899.5

Open-set & adaptation

Recognizing unknown generators, and adapting to new ones

Trained on 17 of OpenFake's 27 generators, the classifier can flag sources it has never seen, and quickly incorporate new ones once labels arrive.

Flagging unknown generators

Images from the 10 held-out sources are flagged as unknown by thresholding the classifier's own confidence — no calibration set, no auxiliary outlier data.

0.875
AU-OSCR, open-set rejection (± 0.037, five draws)

Few-shot adaptation to new generators

The backbone already learns a general fingerprint space, so a generator it has never seen is admitted by freezing that backbone and fitting only a linear head — no retraining. Even a single labeled image per class beats the published baselines, and the head fit takes under four seconds.

Table 3 — Few-shot attribution on nine unseen GenImage generators, N labeled images per class, LIDA's protocol.
Method1-shot10-shot
ResNet17.421.4
DIRE14.317.2
ESSP17.022.4
LIDA40.454.0
Ours (OpenFake backbone)47.7 ± 3.572.0 ± 0.8
Ours (DRAGON backbone)52.0 ± 3.672.4 ± 0.9

Universal feature extractor

Features that generalize beyond attribution

Trained only to classify known generators, the network's features generalize beyond that task: they organize sources it has never seen, and recover the lineage of the ones it has, with no fine-tuning, labels, or calibration.

UMAP of penultimate features for ten generators never seen in training. With no target count, density-based clustering auto-estimates seven clusters against the true 10 sources (ARI 0.64, NMI 0.84); the only collapses are within a shared model family.

The same features, clustered on 27 known OpenFake generators with no lineage labels at all. Architecturally related models land next to each other — the network organizes by structure, not just by which generator produced an image.

BibTeX

@misc{livne2026rpa,
  title         = {Scalable Black-Box Model Attribution for Images},
  author        = {Asaf Livne and Amir Jevnisek and Shai Avidan},
  year          = {2026},
  eprint        = {2608.15652},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2608.15652}
}

Preprint on arXiv; conference citation will be updated on acceptance.