We’ve launched our most advanced model yet: Pangram 4! Read about it here.
We are excited to announce a research preview of an entirely new addition to the Pangram model family: Pangram Image.
Pangram Image is a model for detecting AI-generated images. It detects image outputs from all of the major AI image providers including OpenAI's GPT Image, Gemini's Nano Banana, Midjourney, FLUX, and Grok Imagine, and some AI video providers, including Kling, Seedance, Veo, and Wan.
Pangram Image far exceeds the performance of state-of-the-art academic and commercial AI image detection methods. On an internal benchmark of 5 commercial detectors, Pangram Image achieves 99.5% accuracy, compared to the next closest competitor at 98%. On external benchmarks, such as the NTIRE 2026 AI Image Challenge, Synthbuster, and Mirage, Pangram Image consistently achieves over 99% AUROC and beats competitors by several percentage points.
We understand that from the perspective of artists and creators, false AI accusations of genuinely authentic art cause serious harm, and we took special care to prevent this from occurring with our model. On a subset of 2,000 images in WikiArt, we accurately classified every single image as human. On ReLAION, a diverse dataset of pre-2022 images from the internet, including photos, web graphics, CGI, and more, Pangram Image demonstrates a false positive rate of 0.16%, approximately 1 in 625.
As with our flagship text model, we are openly publishing our methodology and engaging with the research community. Pangram Image is available in the web dashboard for all free and paid Pangram users, and will be coming soon to our Chrome extension and other integrations.
While we are incredibly proud of this model, this is not the final form of Pangram Image, and we will be continuing to iterate on the model and drive up the accuracy in the coming weeks. We believe that we’ve already made great progress, and we are excited to share this early model publicly. During the research preview, access to the Pangram Image API is by invitation only. Please reach out if you are interested in partnering with us as we rapidly iterate on this model.
AI-generated images have gotten incredibly realistic. In the past, you could count fingers or look for malformed text as a sign of AI-generated images. Today, AI images follow instructions and are more realistic than ever.
Additionally, AI-generated visuals are increasingly common to find. Online, you can find AI visuals everywhere from Fruit Love Island TikToks to AI-generated Iran war videos. We hear stories of people using AI to add bugs into a photo of their bowl to get a refund on their delivery. We even see obviously AI-generated visuals printed on menus and signs.
Just like text, not all use cases of AI imagery are harmful. But there are many scenarios where we really want to know. In the recent months, there was an image of Mitch McConnell in the hospital that people speculated was AI (it wasn't). There were AI-generated photos of Taylor Swift's wedding circulating on social media. An AI-generated image was even a finalist in a major photography competition.
Unlike text, there are major efforts to ensure that the provenance of images and videos can be tracked. Efforts like SynthID and C2PA metadata and invisible watermarks to images, identifying the provenance chain of both real and AI images. Unfortunately, watermarks are trivially bypassed by tools that reverse-engineer watermarks and strip metadata. Additionally, while OpenAI and Google both implement SynthID, other AI image providers do not, meaning that relying on watermarks alone will never allow one to be fully confident in the provenance of an image that lacks metadata.
We trained our image model to detect images from a wide range of AI image generators including OpenAI's GPT Image, Gemini's Nano Banana, FLUX, Midjourney, and Grok Imagine. This model can also detect frames from AI video generators including Kling, Seedance, Veo, and Wan.
In most cases, AI edited images will be flagged by our model as fully AI, due to the way modern image generation tools regenerate the entire image even when applying small edits.
We can also detect mixed images with our heatmap feature, which shows you which parts of the image were flagged as AI. Internally, we found this feature useful for testing screenshots or real world photos that include some AI within the scene, but are not fully AI generated. We feel that sharing this heatmap information can help users understand our model better, so we are exposing this heatmap directly in the dashboard.
For this initial research preview, some classes of images are still out of scope. We are currently unable to process images below 512 by 512 pixels, due to the significantly reduced amount of information available to our model in these smaller images. We are also unable to confidently detect deepfakes and face swaps, as we focused on more common consumer image generators in this initial release.
For safety reasons, we are also currently unable to process content flagged by our moderation systems for containing NSFW imagery. We will be able to remove this restriction in the future as we build more safety controls and exit our research preview.
Our model builds upon the industry standard DINOv3 architecture. DINOv3 is a powerful general purpose computer vision representation model, which represents images as vectors, also known as embeddings or features: lists of numbers that encode all of the visual interest in an image in a compressed format.
These embeddings are powerful enough that they contain all the information necessary to solve a variety of computer vision tasks, such as classification, depth estimation, or segmentation, to an exceptional degree of accuracy. Our hypothesis was that given the general purpose nature of these embeddings, and their strength at both understanding local and global properties of images, that they would be a natural starting point for AI image detection.
DINOv3 extracts general-purpose image features that support classification, retrieval, PCA, detection, segmentation, and depth estimation
Image Courtesy: Meta Research Blog
What makes these features so powerful is an advancement of a long line of research called self-supervised learning. These features are learned by a model learning general properties of images on large-scale datasets, rather than being tuned for a specific task. DINO in particular uses a technique called self-distillation, where a “student” neural network must match the output of a “teacher” neural network across differing views and augmentations of an image. The patterns and world knowledge that DINO acquires over the course of its training make it much more powerful than if we randomly initialized a neural network and began learning the AI image detection task from scratch.
We find however, that AI detection is not just an ordinary downstream task. Rather than the typical procedure of simply training a classification head on top of frozen DINO features, we instead use continuation training to fully finetune the backbone along with the AI detection classification head: allowing the DINO features to evolve as our model gains understanding of how human and AI images differ over the course of our second stage of training.
The quality, diversity, and scale of data is vital to any foundation model, and Pangram Image is no exception. We found that the specific data composition of both human and AI-generated imagery had the largest impact on the final accuracy of the model compared to anything else we tried.
We have written extensively about our system for generating synthetic data for our AI text detector, which we call synthetic mirroring. The synthetic mirroring system creates diverse AI examples by prompting AI models to create text that looks very similar, but not exactly identical, to human text. This system grounds each AI example in real topics, styles, genres, and specific pieces of information, and prevents the model from overfitting to differences between pairs of human and AI text that have nothing to do with AI writing style. In fact, the AI dataset for training the text detector is purely synthetic: meaning we generate all of the data ourselves using LLM APIs.
For Pangram Image, we still make use of the synthetic mirroring technique. We use a vision-language model (VLM), a language model that can also take images as inputs, to first generate an extremely detailed caption of a real-world image. We then feed that caption back into an image generation model as a prompt. Sometimes, the images look eerily similar!
A photo of our team and its AI synthetic mirror
A photo of our team and its AI synthetic mirror
One of the challenges that we faced early on when developing Pangram Image was that the types of images our synthetic mirrors produced did not match the distribution of real-world AI images that people were posting online. We realized that our synthetic mirroring technique was insufficient to simulate the challenging distribution of AI images that can be found online “in the wild.” When creating AI generated images, people often crop and edit them in ways that are challenging to simulate.
Our solution was ultimately to rely much more heavily on real world AI images sourced from the Internet, in addition to the images that we synthetically generated in house. We built our real-world datasets in accordance with ethical scraping principles to ensure that we are respectful of the work of creators and do not inadvertently capture personal data or data that the creators would not consent to us using. In addition, we layer strong augmentations during our training process to ensure we capture the ways that images are distorted, compressed, and edited as they’re shared online.
Original image before augmentation
The same image after augmentation — cropped, downscaled, and compressed
Original image before augmentation
The same image after augmentation — cropped, downscaled, and compressed
Original Augmented
We also rely on a new system, composite training, to deliver high quality heatmaps to assist in interpretability of mixed AI/Human imagery. In composite training, we layer AI and Human images to train our model to understand both local and global features. This allows us to present heatmaps in our model, which we find improves user confidence significantly and allows for scanning real-world objects containing AI imagery.
Composite training layers AI and human images so the model learns both local and global features, producing interpretable heatmaps
Robust evaluations are key to our strategy across all of our products. For Pangram Image, we created a large number of evaluation sets to measure our performance across diverse domains. Ethical data sourcing is extremely important to our mission and our values as a company: as part of this commitment, we worked with several people inside and outside of Pangram to source images donated by their owners in support of advancing AI image detection.
We created an internal benchmark of 1,130 example images to compare a variety of commercial AI image detectors in a cost-effective manner. We sourced 500 examples of known human images, including 100 images from our CTO Bradley’s personal iPhone camera roll to ensure that some of the benchmark had never been seen by any of the AI image detection models..
We then asked Claude to create 100 prompts that would reflect real-world use cases of AI image generation, and fed these prompts to 5 frontier AI image generation models: Flux 2 Pro, Bytedance Seedream, Gemini 3.1 Flash Image (Nano Banana), GPT Image, and Riverflow v2. Finally, we sourced 30 images from Midjourney’s website (since they don’t have an API).
We report results on both clean data, as well as augmented data, which involves applying a downscale to 1024x1024, and converting the image to JPEG format with quality 50. This is a higher degree of image compression than most real-world images, so we feel it provides a good measure of quality in worst-case scenarios.
| Detector | Accuracy | FPR | FNR |
|---|---|---|---|
| Pangram Image | 100.00% | 0.00% | 0.00% |
| SightEngine | 99.13% | 0.40% | 1.32% |
| Resemble AI | 98.35% | 2.40% | 0.94% |
| Winston AI | 94.27% | 10.60% | 1.13% |
| AI or Not | 87.67% | 14.40% | 10.38% |
| Detector | Accuracy | FPR | FNR |
|---|---|---|---|
| Pangram Image | 99.03% | 0.40% | 1.51% |
| SightEngine | 97.57% | 0.40% | 4.34% |
| Winston AI | 95.83% | 3.60% | 4.72% |
| Resemble AI | 90.68% | 1.60% | 16.60% |
| AI or Not | 84.17% | 10.20% | 21.13% |
To understand our FPR at scale, we evaluated on pre-2022 images from the ReLAION dataset, sourced from Common Crawl, which provides a highly diverse set of images from all over the web. We filter to pre-2022 to avoid data contamination from image generation models.
| Eval | Images | Clean accuracy / FPR |
|---|---|---|
| ReLAION pre-2022 | 10,000 human | 99.84% / 0.16% |
To investigate whether we also can detect AI generated video content, we also created an internal dataset of video outputs from frontier video generation models. These videos were seeded with real human photos from our dataset.
These video frames were labeled with a Gemini VLM. We report our results across 9 video models: Google Veo 3.1 and Veo3.1 Fast, Wan2.7, Grok Image Video, Kling Video O1 and 3.0 Standard, Seedance 2.0 and 2.0 Fast.
| Video model | N | Clean detected | Clean accuracy | Aug detected | Aug accuracy |
|---|---|---|---|---|---|
| Grok Imagine Video | 276 | 275 | 99.64% | 275 | 99.64% |
| Kling 3.0 Pro | 140 | 139 | 99.29% | 137 | 97.86% |
| Kling 3.0 Standard | 260 | 250 | 96.15% | 245 | 94.23% |
| Kling Video O1 | 120 | 114 | 95.00% | 102 | 85.00% |
| Seedance 2.0 | 140 | 140 | 100.00% | 138 | 98.57% |
| Seedance 2.0 Fast | 216 | 216 | 100.00% | 213 | 98.61% |
| Veo 3.1 | 100 | 99 | 99.00% | 93 | 93.00% |
| Veo 3.1 Fast | 276 | 276 | 100.00% | 274 | 99.28% |
| Wan 2.7 | 220 | 220 | 100.00% | 220 | 100.00% |
| Overall | 1,748 | 1,729 | 98.91% | 1,697 | 97.08% |
One source for these evals was our new @PangramData account. We would like to thank the community for assisting us in understanding real-world AI usage and taking the time to share these findings with us. After manually cleaning the data to remove images that were either out of scope or clearly not AI, we found our model disagreed with the image submissions on only 2 out of the 43 images we received. This is an incredibly helpful dataset for us, as it gives us a strong understanding of the types of images that our users think are AI.
False negative disagreements:
A false-negative disagreement between our model and the community submission
A false-negative disagreement between our model and the community submission
In addition to our internal eval sets, we also compare our results on relevant third-party benchmarks. Pangram Image achieves state of the art performance compared to previous approaches. Different benchmarks report different specific metrics, such as macro accuracy, mAP, or AUROC, so we compute the metrics according to the specific paper and report the closest comparison as reported.
| Benchmark | Pangram result | FPR | FNR | Previous best | Previous best result |
|---|---|---|---|---|---|
| Community Forensics Eval [1] | 97.29% macro accuracy; 99.70% mAP | 0.413% | 6.355% | Ours-384 | 89.3%; 98.7% |
| Synthbuster + RAISE-1K [2] | 98.49% macro accuracy; 99.96% mAP | 0.100% | 2.911% | B-Free | 94.9%; 98.8% |
| Mirage-Test [3] | 99.66% domain-macro accuracy; 99.995% mAP | 0.056% | 0.618% | OmniAID-Mirage | 88.39%; 96.81% |
| AIGenImages2026 [4] | 99.463% accuracy; 99.943% AUROC | 0.537% | 0.537% | RINE | 92.13%; 97.75% |
| NTIRE 2026 Open Test [5] | 99.999% AUROC | 0.500% | 0.000% | MICV | 99.78% AUROC |
One of the best sources that we have for community judgement on whether real world images are AI or not is the subreddit /r/IsThisAI. This makes it a good representative case study for our model’s alignment with human preference, and allows us to discover failure modes in our model. For this case study, we collected a set of 100 posts from the subreddit. Of those, 27 had no clear consensus from the subreddit and as a result were filtered out of our evaluation set. The consensus was judged by a human annotator.
Of the remaining 73 images, 48 were voted as AI by the community, while the other 25 were voted as human. Pangram Image correctly agreed with the reddit decision on all 25 human samples and on 40 out of the 45 of the AI samples as AI. This aligns with our policy to calibrate our models to minimize the chances of false positives occurring.
While we think this is a great model, and we're excited about the future, we know there is a lot more to do. On the research side, we know that performance can be improved on adversarial cases, such as when watermarks are removed or AI images are edited after generation. Additionally, we plan on continuing to drive down the false positive rate to a level closer to that of our text model.
On the product side, expect to see improvements over time. The X bot will soon be able to check for AI images, and the dashboard will be able to accept videos as input. Ultimately, we will be bringing Pangram Image to the Chrome extension, so all Pangram users will get proactive AI labels on images in their social feeds.
1: Park et al., “Community Forensics: Using Thousands of Generators to Train Fake Image Detectors,” CVPR 2025. https://huggingface.co/datasets/OwensLab/CommunityForensics-Eval/blob/main/README.md 2: Qin et al., “Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes,” CVPR 2026, Table 1. https://openaccess.thecvf.com/content/CVPR2026/papers/Qin_Scaling_Up_AI-Generated_Image_Detection_with_Generator-Aware_Prototypes_CVPR_2026_paper.pdf 3: Guo et al., “OmniAID: Decoupling Semantic and Artifacts for Universal AI-Generated Image Detection in the Wild,” https://arxiv.org/pdf/2511.08423 4: Pantsios et al., “Automated In-the-Wild Data Collection for Continual AI Generated Image Detection,” https://arxiv.org/pdf/2605.02567 5: Gushchin et al., “NTIRE 2026 Challenge on Robust AI-Generated Image Detection in the Wild,” https://arxiv.org/pdf/2604.11487
We'd like to thank Artem Frenk for research contributions, Ben Glickenhaus for engineering, Lu Lyu and Fanyi Pan for design and UI, and Tianxu Zhou for Chrome extension integration.

Rachel Stajduhar is a Research Scientist Intern at Pangram. She is graduating from the University of New Hampshire this fall with a degree in Computer Science. In her free time, she loves film photography.

Bradley is an AI researcher and expert in building deep learning products in industry. He recently led the deep learning research group at Absci, a generative AI drug discovery company, and previously was a member of the core computer vision team at Tesla Autopilot.






