Reactive moderation starts with discovery. We measured how often sampled clips contain a searchable person or production identifier and how many distinct people appear in uploader catalogues. The results expose weaknesses in those two routes; they do not measure consent, authorship, or every way an item might be found.
The premise no one measures
Platform reports and DMCA requests usually start with a specific URL located by someone. Statutory regimes differ (the EU Digital Services Act, for example, lets any individual or entity submit a notice rather than limiting notices to victims), but their reactive mechanisms still require someone to locate the material. The research literature establishes the harm of image-based sexual abuse; it does not directly estimate how often platform metadata supplies a searchable person or production identifier.
So we asked two empirical questions. Can the depicted person find a given clip at all? And how is imagery organized across upload accounts: are accounts centered on one person, or are they collections of many? We answered the first by classifying a sample of clips for any searchable identifier, and the second by clustering uploader catalogs by biometric face identity.
A note on framing before the numbers. This is a study of a privacy failure, of how little of this imagery ordinary search can surface for the person in it. It is not a study of copyright or piracy. We treat intimate imagery as a single category and make no claim about origin or consent; neither is observable at scale, and our conclusions do not depend on them. This is not an estimate of NCII prevalence; it measures discoverability-related properties of intimate imagery in the sampled platform content.
Finding 1: about half lack a searchable person or production identifier
We sampled 3,950 clips across four platforms: uniform random draws on three of them, and an uploader walk on erome, which exposes no full catalog to draw from. Each clip was classified for whether its metadata contained a personal name or alias, or a searchable production identifier such as a studio, series, scene, film, or catalog key. A one-word alias counts when it is presented as the depicted person's name; a generic title does not become an identifier merely because its exact phrase can be searched. This metadata proxy does not establish that the depicted person knows the term, that the term is sufficiently specific, or that a search engine retrieves the clip.
The per-site rates are the primary result. Across the three uniformly sampled platforms, 48.1% of 3,000 clips have no searchable person or production identifier. Adding erome's non-uniform uploader-walk sample gives 56.1% in the observed four-platform sample across 3,950 clips. Under a deliberately broad assumption that every verified-creator upload is searchable the latter falls to 54.2%; extending that assumption to studios gives 53.0%. Account status does not prove who is depicted, so those figures are sensitivity bounds rather than observed self-post rates. The historical classification used separate Claude Haiku and Sonnet runs, with Claude Opus adjudicating disagreements. We then completed a locked, model-blind, reviewer-led audit of 400 items and a recorded exact-model cross-provider pass over all 3,950 records with GPT-5.6 Sol.
| Platform | Sample | No identifier |
|---|---|---|
| eporner | 1,000 | 37.4% |
| spankbang | 1,000 | 52.9% |
| tnaflix | 1,000 | 54.1% |
| erome | 950 | 81.2% |
| Uniform draws (3 sites) | 3,000 | 48.1% |
| All four samples | 3,950 | 56.1% |
The per-site pattern is a gradient: eporner is lowest, while erome's uploader-walk sample is highest at 81%. This is consistent with professionally catalogued material carrying names and production codes more often, but the study does not test that explanation causally. Neither pooled value is an account- or catalogue-weighted platform prevalence estimate: 48.1% holds the design to uniform clip draws but excludes erome, while 56.1% includes all four measured platforms but is raised by erome's uploader-walk sample. The latter has a descriptive Wilson interval of 54.5%–57.6% within the observed sample; this is not a platform-population sampling interval. We therefore say about half, not that a general majority has been established.
Finding 2: encountered uploader accounts depict many people
The account frame came from Finding 1. Deduplicating uploaders in the clip sample produced 2,804 accounts: 2,497 labeled as ordinary users, 192 as studios, and 115 as verified creators. The ordinary-user accounts covered 3,557 of the sampled clips. We crawled those catalogs and clustered face embeddings by identity to estimate how many distinct people each account depicts.
Among encountered accounts crawled to at least ten posts, the study reports that 90% depict ten or more biometrically distinct people (a descriptive conditional binomial interval of 88.6% to 91.3% on 1,943 accounts); on the three platforms with the deepest crawls the reported figure is 94%. Those intervals do not include uncertainty from the non-uniform account frame or clustering instrument.
| Platform | Median distinct | Diversity ratio | Depict ≥10 |
|---|---|---|---|
| eporner | 81 | 0.53 | 95% |
| tnaflix | 22 | 0.45 | 96% |
| erome | 80 | 0.53 | 91% |
| spankbang | 5 | 0.50 | 74% |
| Pooled (4 sites) | 0.49 | 90% |
We are careful about what this measures: depiction, not authorship. The count says how many different people appear, not who runs an account or why. Clustering was within accounts, so it does not say how many accounts any one person appears in. Within the encountered set, the observed concentration is inconsistent with a catalogue consisting predominantly of material featuring the account operator alone and is compatible with an account-level collection depicting many different people. The study does not establish how that material was obtained, uploaded, or distributed.
The discovery gap
Together, the findings identify weaknesses in two ordinary routes. Finding 1 weakens metadata search: across the historical pipeline, expanded audit, and recorded exact-model cross-provider pass, the observed shares range from 40.5% to 48.1% on the uniform three and from 49.2% to 56.1% across all four samples. Finding 2 weakens monitoring one natural uploader page: within the encountered account set, many catalogs reportedly depict many different people.
Discovery deserves measurement as its own stage. A fast takedown process cannot act on a specific item that nobody has located.
The narrowness matters. An acquaintance can recognize someone; a depicted person who already has the source image can use reverse-image search; consumer face search or proactive platform systems may surface material. We did not measure those routes. The study shows why a notice-only workflow can miss material before takedown begins; it does not estimate the share that is never found by any means.
What the study does not establish about face search
We did not test consumer face-search coverage on these platforms, so the study cannot say that existing products fail or succeed. PimEyes' own guide says it checks publicly accessible pages, including pages with explicit content, and that choosing Safe Search hides explicit results. Coverage and recall remain open empirical questions.
Nor does this paper declare a biometric self-search product lawful. Biometric recognition can engage special-category data rules; UK ICO guidance, for example, requires an Article 9 condition as well as an Article 6 lawful basis. Consent, self-search, authorized-agent access, security controls, retention, jurisdiction, and governance all need separate analysis. The measurement motivates that work; it does not settle it.
What can be verified today
The complete study record exists, though it is split across restricted local and GPU-host storage. It includes the 3,950 per-clip records and historical classifier outputs, both initial 100-item audit files, the corrected expanded 400-item audit and weighted analysis, the TNAFlix repair ledger and aggregate receipts, the complete corrected GPT-5.6 Sol robustness output and provenance receipts, the 2,497 account-level Finding 2 rows and 742 clean spankbang re-crawls, the repaired 50-account validation record, its archived predecessor, the completed 469-pair de-identified verdict output and correction receipts, analysis scripts, and calibration summaries. One historical provenance gap remains: the exact Sonnet and Opus identifiers were not written to the outputs, logs, or shell history; the new GPT run records its exact requested and returned model plus prompt, input, and output hashes.
Replication without publishing identities
Existence does not imply that every record should be public. Source metadata can contain explicit names or aliases linked to adult content; account and validation records can contain uploader handles, URLs, crop paths, and other linkage material. Publishing those files would create a new exposure while measuring an existing one.
The versioned public verification package is available for download. It contains the analysis code, prompts and codebook, sampling notes, aggregate tables, calibration summary, safe cross-provider aggregates and provenance, the expanded audit's collapsed findable/unfindable/ambiguous labels plus the non-identifying design fields required to reproduce its weighted prevalence, confidence intervals, and reviewer–pipeline agreement, the de-identified correction and completion receipts, and the completed candidate-pair verdict output under de-identified account keys. Its README gives a clean-room verification command, its checksum list hashes every file, and its license covers reuse of the code, de-identified data, and documentation. The audit interface also recorded whether the reviewer chose name, production identifier, or both, but those reason categories were not applied consistently in the first 100 items. They are neither analyzed nor released; all three reliably collapse to “findable.” Raw metadata, names, aliases, URLs, uploader identifiers, face crops, embeddings, review-image paths, and reversible linkage keys will not be published. Additional derived data would be considered only under controlled access where disclosure risk can be adequately managed. An identical repository copy is deposited with Figshare under DOI 10.6084/m9.figshare.33263424.v1.
Release integrity. Version 1.4, released 12 August 2026. PDF SHA-256: 44771e3c31878b1958748d2635ae140faedb058761b17875e110b1ee9fddf821. Verification ZIP SHA-256: 7661ddb2a3c3b30768da524123159620052ca797553817b81def84c5e8f53ff2.
Citation. QayemDigital Research, “Findable: A Biometric Measurement of Intimate Imagery Beyond the Reach of Ordinary Search,” version 1.4, 12 August 2026. This page is the version of record. An archived copy of the paper and verification package is deposited with Zenodo under DOI 10.5281/zenodo.21895855.
Ethics
This study measures a privacy harm under strict minimization. Every public figure is aggregate, and the paper identifies, ranks, or links no sampled individual to an uploader or real-world identity. No intimate scenes or full frames are published. Calibration and quality control used restricted face-only crops and biometric templates; performer labels were used only to evaluate the threshold, and study accounts were not cross-searched against named identities. Those sensitive artifacts are not shared. They will be deleted within 180 days of the verification package release, and no later than 30 June 2027, with the deletion recorded in a signed receipt appended to the package.
QayemDigital Research is independent and has no institutional review board. In place of one, the study ran under a written protocol with definitions locked before analysis, and the paper structures its ethics reasoning around the Menlo Report principles so the reasoning itself can be examined. Crawling used publicly served pages only (no accounts, no logins, no paywall or age-verification bypass) with paced requests that accepted a shallower crawl where a platform rate-limited. For the biometric processing, the paper identifies the research-purposes condition with safeguards as a candidate Article 9 basis; it does not claim that this supplies the separate Article 6 analysis or a formal data-protection impact assessment, and it offers no external legal opinion. A documented impact assessment and jurisdiction-specific legal review are required before future biometric processing or any product deployment. The analysis does not rest on a claim that depicted persons made their data public, which would assume the self-posting the study does not assume.
Exposure minimization also shaped the validation design. One trained reviewer saw only the metadata and face-only material needed for the audits. Adding reviewers would have widened access to explicit identifiers and biometric crops derived from sensitive content. This choice prevented measurement of inter-rater reliability, so we report it as both an ethical safeguard and a methodological limitation, not as evidence of accuracy. It is a standing position rather than a scheduling gap: no additional person will be shown the study material, including titles and tags. The expanded metadata audit was reviewer-led rather than human-only: for some difficult translations and codebook-edge cases, the reviewer consulted an AI assistant or dictionary without revealing historical model verdicts, strata, platform labels, or linkage keys. Source pages were consulted in a small number of opaque abbreviation or account-status cases. Conservatively treating all four source-assisted findable records as unfindable moves the corrected audit from 52.6% to 53.1% overall and from 43.9% to 44.6% on the uniform three. These qualifications are disclosed rather than treated as independent human-only validation.
Depending on the platform, the historical pipeline finds no searchable person or production identifier in 37% to 81% of sampled clips. Its pooled result is 48.1% across the three uniform clip samples and 56.1% in the observed four-platform sample after adding erome's non-uniform uploader walk; the expanded audit and GPT-5.6 Sol checks find somewhat lower but still substantial gaps. None is presented as a general majority estimate. For assessable ordinary-user uploader accounts encountered through those clips, the study reports that 90% depict ten or more biometrically distinct people. The first result qualifies metadata search; the second qualifies account-by-account browsing. Both point to discovery as a stage worth measuring separately from removal.
Read the full paper, with methods and limitations (PDF) → Download the verification package (zip) →