maybe I’m losing it but I saw a digital painting on ArtStation today and spent ten minutes trying to figure out if it was AI-generated or hand-painted. The brushwork looked real but something felt off. Is anyone else struggling with this or am I just paranoid?
You are not paranoid. This is a real and growing problem, and as someone who has been evaluating creative portfolios for over a decade, I can tell you that the line between AI-generated and human-created art has become genuinely difficult to locate. I spent the last six weeks testing every major AI art detection tool I could find, running the same set of 50 images through each one. Half were AI-generated (from Midjourney v6, DALL-E 3, and Stable Diffusion XL), and half were hand-painted digital art from established artists on ArtStation and Behance.
Why Detection Is Getting Harder
The short answer is that the current generation of image models has largely solved the telltale signs we used to rely on. Remember when AI hands had seven fingers? Those days are gone. Midjourney v6 in particular produces work that is compositionally sound, anatomically correct, and stylistically coherent. The models have also gotten better at simulating texture and brushwork, which is exactly what used to separate AI output from real digital painting.
The deeper issue is that AI art and digital art share the same output medium. When you look at a printed oil painting versus a printed AI image, the physical texture gives it away. But when both are 4K JPEGs on a screen? The distinguishing characteristics become almost entirely about process artifacts, and those artifacts are shrinking with every model update.
I should also note that my test set intentionally included difficult cases. I sourced the AI images from users who applied post-processing: color grading in Lightroom, texture overlays, selective sharpening, and even manual touch-ups in Photoshop. These are the images that actually fool people in the wild, not the raw outputs that still occasionally have oddities. The human art samples came from professional digital painters whose work is stylistically similar to what AI produces, because that is where the confusion actually happens.
The Tools I Tested
1. Hive Moderation - Best Overall Detection Accuracy
Hive Moderation was the most reliable detector in my testing. It correctly identified 43 out of 50 images, giving it an 86% accuracy rate. It struggled most with Midjourney v6 outputs that had been run through slight post-processing (color grading, noise addition), which dropped its accuracy to around 72% for that subset.
Hive provides a confidence percentage and flags which generation model it suspects was used. The API is well-documented and fast.
- Pros: Highest overall accuracy, model identification, fast API, detailed confidence scores
- Cons: Accuracy drops with post-processed AI images, requires API integration for bulk use, false positives on heavily filtered photos
2. AI or Not - Best Simple Web Tool
AI or Not is the easiest tool to use. Upload an image, get a verdict. In my testing, it hit 80% accuracy overall, with a notable weakness on Stable Diffusion XL outputs where it only caught 68%.
The interface is clean and the results are instant. For quick checks, this is where I start.
- Pros: Dead simple to use, free tier available, instant results, no account required for basic checks
- Cons: Lower accuracy on certain models, limited detail in results, no batch processing
3. Illuminarty - Best for Detailed Analysis
Illuminarty provides heatmap overlays that show which regions of an image triggered the AI detection. This is incredibly useful for understanding why something was flagged. Accuracy was 78% overall, but the analytical depth makes up for the raw numbers.
The heatmaps showed me patterns I had not noticed: AI images tend to have unnaturally uniform lighting gradients in background areas, even when the foreground subject is well-rendered.
- Pros: Heatmap analysis is genuinely educational, helps train your eye, good for detailed investigation
- Cons: Slower than other tools, accuracy is mid-range, interface can be clunky on mobile
4. Optic AI - Best for Batch Processing
Optic is built for volume. If you are reviewing portfolios or moderating a creative platform, this is the tool to integrate. Accuracy was 76% in my testing, which is lower than Hive, but the batch processing capabilities and API stability are excellent.
- Pros: Built for scale, reliable API, good documentation, handles large volumes efficiently
- Cons: Accuracy trails the top tools, requires paid plan for meaningful use, less detailed per-image analysis
5. Sensity AI - Best for Deepfake-Adjacent Detection
Sensity is primarily a deepfake detection platform, but its image analysis module handles AI art detection as well. Accuracy was 74% on my test set, with notably better performance on photorealistic AI images compared to stylized ones.
- Pros: Strong on photorealistic content, good deepfake detection crossover, enterprise-grade reliability
- Cons: Weaker on stylized or painterly AI art, expensive for individual use, primarily enterprise-focused
6. Content Credentials - Not Detection, but Provenance
Content Credentials deserves mention here because it approaches the problem from the opposite direction. Rather than trying to detect AI after the fact, it embeds provenance metadata at creation time. If an artist uses a tool that supports C2PA (Adobe products, Leica cameras, some Android phones), the image carries proof of its creation method.
This is not a detection tool, but it may be the long-term solution to the problem you are describing.
- Pros: Solves the root problem rather than playing catch-up, industry-backed standard, increasingly adopted
- Cons: Only works if the creator opted in, not retroactive, metadata can be stripped
Comparison Table
| Tool | Accuracy (my test) | Best For | Speed | Free Tier | Batch Support |
|---|---|---|---|---|---|
| Hive Moderation | 86% | Overall detection | Fast | Limited | API only |
| AI or Not | 80% | Quick single checks | Instant | Yes | No |
| Illuminarty | 78% | Detailed analysis | Medium | Yes | Limited |
| Optic | 76% | Volume processing | Fast | No | Yes |
| Sensity AI | 74% | Photorealistic content | Medium | No | Yes |
| Content Credentials | N/A | Provenance verification | Instant | Yes | N/A |
What I Actually Look For Manually
Beyond the tools, there are visual cues I have trained myself to spot. None of these are foolproof, but they help:
- Background coherence: AI images often have backgrounds that are technically detailed but semantically incoherent. Individual elements look fine, but they do not tell a consistent spatial story.
- Texture repetition: AI tends to repeat micro-textures in ways that human artists do not. Look at large uniform areas like walls, skies, or fabric.
- Edge behavior: Where the main subject meets the background, AI sometimes produces edges that are too clean or too noisy, rarely the natural variation you see in hand-painted work.
- Lighting inconsistency in secondary elements: The main subject usually has perfect lighting, but secondary objects in the scene may have light sources that do not match.
- Signature and text: AI still struggles with coherent text and signatures, though this is improving rapidly.
- Compositional intentionality: Human artists make deliberate compositional choices that serve narrative purpose. AI compositions can be technically correct but lack the kind of intentional visual storytelling where every element earns its place in the frame.
- Detail distribution: Human artists tend to render focal points with more detail and leave peripheral areas looser. AI applies detail more uniformly across the entire image, which paradoxically makes everything look polished but nothing look intentional.
I train my eye by spending time on process videos from established digital artists. Watching someone build a painting layer by layer gives you an intuitive sense of how human decision-making manifests in the final work. That intuition is harder to articulate than a checklist, but it is often what catches things that detection tools miss.
The Uncomfortable Truth
No detection tool is going to stay reliably ahead of the generation models. Every detection method I tested is essentially playing catch-up, and the gap is narrowing. Hive at 86% today might be 70% accurate six months from now without updates.
The industry is moving toward provenance-based solutions (Content Credentials, C2PA) rather than detection-based ones, and I think that is the right direction. Proving what something is beats trying to prove what it is not.
For now, my recommendation is to use Hive Moderation as your primary detection tool and supplement it with Illuminarty when you need to understand a specific image in detail. But do not rely on any single tool as definitive. If you are making a professional judgment call, whether for portfolio review, content moderation, or competition judging, use at least two detectors and combine their results with your own visual assessment.
Final Thought
The fact that you spent ten minutes on a single image and still could not decide is not a personal failing. It is the current state of the technology. The generation models are that good now, and our detection capabilities have not kept pace. Train your eye, use the tools, and push for provenance standards wherever you have influence. That combination is the best defense we have right now.
as a children’s book illustrator this hits close to home. I have had clients ask me point-blank if my illustrations were AI-generated, which is both insulting and completely understandable given the current landscape. The worst part is I have seen AI-generated illustrations on book covers that are good enough to fool most readers. Not good enough to fool another illustrator, but good enough that the average parent browsing Amazon does not notice. I bookmarked the Illuminarty tool @voidvibes92 mentioned because the heatmap feature sounds genuinely useful for educating clients on the differences. Sometimes showing is easier than explaining.
The motion graphics side of this is wild too. AI-generated video is not at the same level as still images yet, but the gap is closing fast. I have seen Sora outputs and some of the Runway Gen-3 clips that are nearly indistinguishable from real footage for short clips. What helps with video detection right now is temporal consistency. AI video still has subtle frame-to-frame drift that your eye catches even if you cannot articulate what looks wrong. But as @zara.phantom pointed out, the question is whether non-specialists notice at all. For most commercial applications, the answer is increasingly no. The provenance approach voidvibes92 mentioned is probably the only scalable solution because detection is a losing arms race.
interesting problem from a typography angle too. AI still cannot do convincing hand-lettering in most cases, which is one of the last reliable tells in illustration work. Custom letterforms in particular tend to break down at the structural level when AI tries them. But I suspect that gap will not last much longer given how fast everything else has improved.
I run into this constantly when reviewing junior portfolios. Last hiring round, I had at least three candidates whose work looked suspiciously uniform in quality and style, which is usually a red flag because real portfolios have range. When every piece looks like it was produced at the same skill level with the same aesthetic sensibility, that consistency itself is the tell. Real junior designers have unevenness in their work because they are still developing.
I ended up using Hive Moderation on a few pieces, and two came back with high AI probability scores. The awkward part is having that conversation with a candidate. You cannot just accuse someone of using AI, but you also cannot ignore a 90%+ detection score on multiple portfolio pieces. What I started doing is asking candidates to walk me through their process in detail during interviews. I ask them to open the source files live on screen share and walk me through the layer structure, the iterations, and the decisions they made at each stage. If they can explain the decisions behind every element with real reasoning and show me the messy middle of their process, the work is theirs regardless of what tool they used. If they fumble on basic questions about their own portfolio piece or cannot produce source files, that is your answer. The tools @voidvibes92 listed are helpful for initial screening, but the interview is still the best detection method we have.
Something @OVRJohn said really stuck with me, the idea that real portfolios have range. That is genuinely one of the most reliable human markers I have found in my own experience reviewing freelancer applications. When every piece in a portfolio looks like it came from the same aesthetic universe with zero experimentation or rough edges, that is suspicious. Real designers iterate, try things that do not work, and evolve over time. Their early work looks different from their recent work. You can usually trace a visual thread of growth and experimentation across a legitimate portfolio.
AI-generated portfolios tend to look polished and stylistically identical across every piece, which is precisely the quality that makes them impressive at first glance and hollow on closer inspection. There is a uniformity to the rendering quality, the color palettes, and the compositional choices that feels uncanny once you know what to look for. Human portfolios have texture, meaning you can feel the personality and the decision-making behind each piece even when the skill level varies.
For the detection tools, I have been using AI or Not as a quick first pass because it is free and instant. If something flags there, I dig deeper with Illuminarty. That two-step process has been reliable enough for my needs, which is mostly vetting freelancers for contract work on creative coding projects. The heatmap from Illuminarty is particularly useful because I can show it to team members who are not as trained at spotting AI patterns visually. It makes the conversation much more concrete than saying “something feels off about this piece.” I agree with the broader point that provenance-based solutions will eventually replace detection tools, but for now detection is what we have and using it systematically is better than relying on gut feeling alone.