Image detector
Last updated 2026-09-24 · heads: tayanch_union_vitl14_v5.joblib, tayanch_union_semantic_clip.joblib, tayanch_edit_head_v2.joblib, inpaint_localizer_v2.pt
A frozen OpenCLIP ViT-L/14 embedding plus 17 camera-physics features, scored by a small head (union v5). Previews and web copies with no camera data go to a CLIP-only sidecar that can only lower the verdict, never raise it; photos judged real pass an edited-photo screen that can only demote to uncertain; the heat map comes from a separate localizer.
Intended use
- Photographs and generated stills a person could mistake for a photograph.
- Triage before a human decision: a verdict is evidence, not proof.
Out of scope
- Screenshots, scans, documents and slides.
- Artwork, illustrations, cartoons and 3-D renders (real or not).
- RAW camera files (not decoded), and photos smaller than a social preview.
- Deciding whether a real photo was retouched: the edit screen flags, it never accuses.
Data and licences
Union pool of about 31,000 images: real photos from COCO, ImageNet, Wikimedia, Food-101, Places365, DOCCI, phone and iPhone originals and a web sample; AI images from 58 generator families (Stable Diffusion 1.4 to 3, SDXL, DALL-E 3, Midjourney v6, DeepFloyd IF, Civitai, BigGAN, ProGAN, StyleGAN3 and the OpenFake frontier sets). Each dataset is used under its own research licence; a licence review for commercial use is open.
Evaluation
| What was measured | Result | n | 95% CI | Date |
| Generators never used in training (head only) | Gemini 97.7%, FLUX.2-Klein 66.0% recall | whole families held out; per-family n not in the tracked record | | 2026-07-17 |
| Test-fold AUC (in-distribution, not a generalisation claim) | 0.9959 | test fold of the 31k pool (n not in the tracked record) | | 2026-07-17 |
| Real photos flagged AI, head only at thr 0.99 | 0.75% | 9 / 1,200 in-domain | 0.40–1.42% | 2026-08-23 |
| Real photos flagged AI, head only, secondary training-pool sources | 0.25% | 2 / 800 | 0.07–0.91% | 2026-08-23 |
| Live API, AI images called AI (development bench) | 94.5% | 52 / 55 | 85.1–98.1% | 2026-09-11 |
| Live API, real items called AI (development bench) | 0% | 0 / 57 | 0–6.3% | 2026-09-11 |
| Edit screen fires on real web copies | 4.7% | 2 / 43 Picsum | 1.3–15.5% | 2026-09-05 |
| Edit screen fires on phone originals (4032 px) | 0% | 0 / 40 | 0–8.8% | 2026-09-05 |
Operating point
AI at p_ai ≥ 0.99, real at ≤ 0.20, uncertain between. Sidecar 0.8255 / 0.1901 on laundered copies. Edit screen 0.4998 (about 2 in 100 untouched photos).
Known failure modes
- Generators unlike anything trained on are caught less often (FLUX.2-Klein 66.0%); faces_ai 0.60 and SD 1.4 0.75 recall at the fit's operating point.
- EXIF-stripped web copies lose camera traces; the physics head alone would read them as AI, so they are routed to the sidecar, which is weaker on unseen generators.
- The development bench shares sources with training; no fully held-out real-source measurement through the served routing exists yet.
- Edited-photo recall on images above 1536 px is unmeasured; against an editor it was not trained on the screen's recall is 0.317.
Video detector
Last updated 2026-09-24 · heads: tayanch_video_frame_clip_v3.joblib, tayanch_video_seq_head.pt, detector_v2_video_ff_dinov2_slim_bml.joblib, detector_v2_t2v_temporal_v2r.joblib
Sampled frames are scored by a frame head over the same CLIP vector as the image model, and a temporal self-attention head reads the sequence; published rules (config/video_rules.json) combine them. The FaceForensics++ face head and the text-to-video head are evidence only: neither can accuse on its own.
Intended use
- Short social clips and uploads: is this footage generated (text-to-video) or camera footage?
Out of scope
- Face-swap / reenactment detection is not claimed: the face head is evidence only.
- Animation, cartoons, games and screen recordings.
- Very short or very low-bitrate clips (a 5.9 s, 360p, 0.4 Mbps real clip was accused).
- Audio: the soundtrack is not judged.
Data and licences
Frame head v3: 1,400 platform-reel AI clips (Kling, Veo 3, Sora 2, Seedance, Runway, Wan 2.2, Higgsfield and generic AI tags) and 1,200 Hugging Face AI clips against 8,500 real reels from 39 social sources plus public real-video sets; labels reviewed on contact sheets. Face head: FaceForensics++ (research licence). Platform clips are used for research evaluation; a licence review for commercial use is open.
Evaluation
| What was measured | Result | n | 95% CI | Date |
| Held-out sources (14 generator + 8 real sources): real clips accused | 0.73% | of 686 real clips | | 2026-09-13 |
| Held-out sources: AI clips read as confidently real | 13.0% | of 2,209 AI clips | | 2026-09-13 |
| Temporal head, held-out AUC | 0.9916 | held-out families and real domains | | 2026-09-11 |
| Served bench of real videos before the face-head fix: accused | 6.4% | 18 / 281 | 4.1–9.9% | 2026-09-23 |
| Same bench, stored results replayed with the face head as evidence only | 0.4% | 1 / 281 | 0.1–2.0% | 2026-09-23 |
Operating point
Rules in config/video_rules.json (frame and attention cut points swept for real accused ≤ 0.5%, real uncertain ≤ 10%, AI read as real ≤ 5% on held-out groups).
Known failure modes
- Reels tagged Higgsfield (51%) and Sora 2 (55%) still read as real too often.
- Real vlogs (16%) and dashcam clips (9%) go to uncertain more often.
- The replayed fix still needs a fresh run on new real footage to confirm.
Audio screen (advisory)
Last updated 2026-09-24 · heads: detector_v2_audio_cube_v2.pt
A six-view spectrogram cube read by a ResNet-18 head with a domain head (speech, music, environment) and a real-sound manifold distance. It has no measured operating point on real-world audio, so every reply is advisory: verdict uncertain, advisory_only true.
Intended use
- Information only: a score and the sound domain, shown next to other evidence.
Out of scope
- Any verdict. The audio screen never calls a file AI or real.
- A verdict on a video soundtrack: the soundtrack row under a video is the same advisory score.
Data and licences
About 30,000 clips: real and synthetic speech, music and environmental sound from public corpora; 24 synthesis systems trained, 8 held out. Each corpus is used under its own licence.
Evaluation
| What was measured | Result | n | 95% CI | Date |
| Cube head v2, held-out (source- and system-disjoint) AUC | 0.898 | | 0.888–0.908 | 2026-09-19 |
| Cube head v2, real sound flagged at the candidate thresholds | 9.3% | | 8.2–10.5% (gate needs ≤ 5%: FAILED) | 2026-09-19 |
| ASVspoof speech head on platform audio | AUC 0.53 | | chance; retired from verdicts | 2026-09-05 |
Operating point
None: uncalibrated, advisory only.
Known failure modes
- Failed its release gate: 9.3% of real sound would be flagged.
- Speech in languages not trained on and music raise false alarms (v1 inverted on music).
- An earlier speech head accused 3 of 5 real clips live on 2026-09-05 before it was made advisory.