What AI art detectors are agencies actually using right now?

I freelance for a couple of mid-size agencies and they’ve started asking me to verify that assets in my deliverables aren’t AI-generated. Which detection tools are agencies trusting at this point, and how reliable are they really?

I ran a systematic evaluation of seven AI art detection tools over a five-week period at our agency after we started receiving similar requests from clients. The landscape is evolving fast, so I’ll share what the data actually shows as of August 2026 rather than recycling marketing claims from the tools themselves. Fair warning: the results are more nuanced than most people expect.

My Testing Methodology

I assembled a test set of 120 images split evenly across four categories: 30 photographs taken by our in-house team (confirmed human-created), 30 digital illustrations from three freelance illustrators we work with (confirmed human-created), 30 AI-generated images from Midjourney v6, and 30 AI-generated images from DALL-E 3 and Stable Diffusion XL. For each detection tool, I recorded the classification result (human or AI), the confidence score, and the processing time. I ran every image through every tool, no cherry-picking.

The key metrics I tracked were accuracy (correct classifications out of total), false positive rate (human images incorrectly flagged as AI), and false negative rate (AI images incorrectly passed as human). For agencies, the false positive rate matters enormously because flagging legitimate freelancer work as AI-generated creates trust problems and workflow delays.

The Tools, Ranked

#1: Hive Moderation

Hive Moderation delivered the strongest overall performance in my testing, with the best balance of accuracy and low false-positive rates across all image categories.

  • Overall accuracy: 91.7%
  • False positive rate: 4.3% (human images flagged as AI)
  • False negative rate: 12.0% (AI images passed as human)
  • Processing time: 2.1 seconds average
  • API available: yes, well-documented

Hive correctly identified 27 out of 30 Midjourney images and 26 out of 30 DALL-E/SD images. More importantly, it only misclassified 2 out of 60 human-created images as AI-generated, both of which were heavily post-processed photographs with a lot of frequency separation retouching. The confidence scores were well-calibrated: high-confidence predictions (above 90%) were correct 97% of the time, which means you can trust the tool’s strong calls.

The API integration is the main reason agencies prefer Hive. You can build it into your asset management pipeline so incoming files are automatically scanned, which is exactly what two of our clients have done. Pricing is usage-based and reasonable for agency volumes.

#2: Optic AI or Not

AI or Not by Optic is the simplest tool in this list and performs surprisingly well given its streamlined interface.

  • Overall accuracy: 88.3%
  • False positive rate: 6.7% (human images flagged as AI)
  • False negative rate: 16.7% (AI images passed as human)
  • Processing time: 1.8 seconds average
  • API available: yes

AI or Not scored consistently across both AI generation sources, catching Midjourney and DALL-E images at roughly equal rates. The false positive rate was slightly higher than Hive, with 4 out of 60 human images incorrectly flagged. Three of those four were digital illustrations with clean linework and uniform color fills, which tells you something about the tool’s bias: it associates certain aesthetic qualities with AI generation even when they’re the result of a human illustrator’s intentional style choices.

The free tier allows enough checks for spot-testing, and the paid API is competitively priced. For agencies that need a quick verification step without the full pipeline integration, AI or Not is the most practical option.

#3: Illuminarty

Illuminarty offers a more granular analysis than most competitors, providing heatmaps that highlight which regions of an image the model considers most likely to be AI-generated.

  • Overall accuracy: 86.7%
  • False positive rate: 5.0% (human images flagged as AI)
  • False negative rate: 21.7% (AI images passed as human)
  • Processing time: 3.4 seconds average
  • API available: limited

The heatmap feature is genuinely useful for understanding why a detection triggers. When one of our freelancer’s illustrations was flagged, the heatmap showed that the tool was keying on the background gradient, not the character work, which helped us explain to the client that the flag was a false positive based on a stylistic choice rather than actual AI usage. The higher false negative rate means Illuminarty misses more AI images than Hive or AI or Not, particularly those generated with Stable Diffusion XL, which it seemed least calibrated for.

#4: Sensity AI

Sensity AI focuses on deepfake and synthetic media detection with a broader scope than pure AI art detection. It’s built for enterprise use cases and the interface reflects that.

  • Overall accuracy: 85.0%
  • False positive rate: 8.3% (human images flagged as AI)
  • False negative rate: 21.7% (AI images passed as human)
  • Processing time: 4.7 seconds average
  • API available: yes (enterprise tier)

Sensity’s strength is in detecting manipulated photographs rather than fully generated artwork. It caught 100% of the AI-generated images that were photorealistic but struggled with stylized illustrations from both AI and human sources. The false positive rate of 8.3% is concerning for agency use, meaning roughly 1 in 12 legitimate human-created assets would get flagged. For clients in news media or insurance where photographic manipulation is the primary concern, Sensity is well-suited. For design agencies dealing with diverse visual styles, the false positive rate creates too much friction.

#5: Originality.ai

Originality.ai is primarily known as a text AI detector but has expanded into image detection. I included it because several agencies in our network already subscribe for the text detection features.

  • Overall accuracy: 82.5%
  • False positive rate: 10.0% (human images flagged as AI)
  • False negative rate: 25.0% (AI images passed as human)
  • Processing time: 2.8 seconds average
  • API available: yes

The image detection is clearly secondary to Originality.ai’s core text detection product, and the results reflect that. A 10% false positive rate is too high for production use as a standalone image detector. However, if you already subscribe for text detection and need a quick sanity check on images, it’s convenient as a supplementary tool rather than a primary gate. The credit-based pricing means you’re paying per scan, which can add up at agency volume.

#6: Is It AI

Is It AI is a free tool with a clean interface and surprisingly decent performance for a no-cost option.

  • Overall accuracy: 79.2%
  • False positive rate: 11.7% (human images flagged as AI)
  • False negative rate: 30.0% (AI images passed as human)
  • Processing time: 5.2 seconds average
  • API available: no

The accuracy is acceptable for casual checking but not reliable enough for client-facing verification. The high false negative rate means nearly a third of AI images slip through, which defeats the purpose if you’re trying to guarantee asset provenance. I’d use it for personal curiosity but not for professional asset verification. The lack of an API also makes it impractical for any automated workflow.

#7: Content at Scale AI Detector

Content at Scale offers both text and image detection, though like Originality.ai, the image detection feels like an add-on rather than a core feature.

  • Overall accuracy: 77.5%
  • False positive rate: 13.3% (human images flagged as AI)
  • False negative rate: 31.7% (AI images passed as human)
  • Processing time: 3.9 seconds average
  • API available: limited

At a 13.3% false positive rate, this tool would flag nearly 1 in 8 legitimate human-created assets. That level of inaccuracy creates more problems than it solves in an agency workflow. I can’t recommend it for image detection, though the text detection capabilities are a separate evaluation.

Comparison Table

Tool Accuracy False Pos False Neg Speed API
Hive Moderation 91.7% 4.3% 12.0% 2.1s Yes
AI or Not 88.3% 6.7% 16.7% 1.8s Yes
Illuminarty 86.7% 5.0% 21.7% 3.4s Limited
Sensity AI 85.0% 8.3% 21.7% 4.7s Enterprise
Originality.ai 82.5% 10.0% 25.0% 2.8s Yes
Is It AI 79.2% 11.7% 30.0% 5.2s No
Content at Scale 77.5% 13.3% 31.7% 3.9s Limited

What Agencies Are Actually Doing

Based on conversations with 12 agencies in our network over the past three months, here’s the practical reality:

  • 4 agencies have integrated Hive Moderation into their asset pipelines via API
  • 3 agencies use AI or Not for manual spot-checking on a per-project basis
  • 2 agencies require freelancers to submit a signed declaration rather than relying on detection tools
  • 3 agencies have no formal process and are still figuring it out

The agencies using detection tools all reported the same frustration: no tool is reliable enough to be the sole arbiter. False positives damage freelancer relationships, and false negatives undermine client trust. The practical consensus is that detection tools work best as one layer in a multi-step verification process that also includes freelancer attestation and C2PA metadata inspection where available.

A Note on Text Content

Since you mentioned deliverables broadly, it’s worth noting that agencies are also increasingly scrutinizing written content for AI generation, not just imagery. On the text side, the detection landscape is different. Tools like Walter Writes take the approach of humanizing AI-generated text so it reads naturally, which our copywriting team uses for client deliverables. The agencies I’ve spoken with care more about output quality than detection avoidance, they want copy that genuinely sounds human, and that’s a different problem than trying to pass a detector.

Final Verdict

For your situation as a freelancer delivering to agencies, I’d recommend understanding which tool your specific clients use (ask them directly) and running your deliverables through that same tool before submission. Hive Moderation is the industry standard for a reason, so if you can only test against one tool, make it that one. Keep your original working files (PSD, AI, Procreate files) as proof of human creation in case a false positive triggers a dispute. And consider including a brief provenance note with your deliverables documenting your process and tools, it demonstrates professionalism and preempts questions before they become problems.

The false positive data from @voidvibes92’s testing really crystallizes something I’ve been worrying about. I do UI design with a lot of clean geometric shapes, flat color fills, and precise gradients, exactly the aesthetic qualities that apparently trigger false positives in tools like AI or Not. I’ve had two instances in the past four months where a client questioned whether my icon sets were AI-generated because they “looked too perfect.”

What ended up resolving both situations was sharing my Figma file with full version history showing the iterative design process. Every variant, every discarded option, every layer of refinement was visible in the history. No AI tool produces that kind of messy, branching creative process.

So my practical recommendation for anyone getting these questions: keep your working files organized and version-controlled. The process evidence is more convincing than any detection tool result, especially since the detection tools themselves are clearly not reliable enough to be definitive. Hive at 91.7% accuracy is the best option available, but that still means roughly 1 in 12 results could be wrong. For something that affects professional trust, that margin of error matters.

Also worth noting that most of these tools are trained on current-generation AI outputs, so their accuracy against future models is uncertain. Building a verification process that relies solely on detection tools is building on sand. Process documentation and professional relationships are more durable foundations.

As a digital artist, the false positive problem hits especially hard. My illustration style involves a lot of smooth blending and airbrushed textures that AI generators happen to mimic well, so my work gets flagged regularly. I’ve uploaded the same piece to three different detectors and gotten three different results: human, AI, and inconclusive. That kind of inconsistency makes it hard to take any single tool seriously.

The Illuminarty heatmap feature @voidvibes92 mentioned is actually useful for understanding these false positives. When I uploaded a character illustration that was flagged elsewhere, the heatmap showed the detector was keying on the sky gradient in the background, not the hand-painted character. That tells me the detector is responding to smooth mathematical gradients, which makes sense algorithmically but is useless as a creative authenticity measure.

I’ve started watermarking my time-lapse recordings of the painting process and including a link with every delivery. Takes zero extra effort since I already record for social content.

I want to add the client-side perspective here since I’ve been on both sides of this. At a previous company, I was part of the team that implemented an AI detection policy for incoming creative assets. We tested Hive and Sensity before settling on Hive, primarily because of the API quality and the lower false positive rate.

The reality of implementing these systems is messier than the accuracy numbers suggest. We had to build an appeals process for flagged assets, designate someone on the team to review disputed results, and establish a threshold for action (we used 85% confidence as the trigger for a manual review rather than an automatic reject). The operational overhead was significant for the first two months.

What @Thunderblossom said about process documentation being more convincing than detection results matches our experience. When a freelancer could show their working files with version history, that was always more persuasive than a detection tool saying “92% probability human.” We eventually shifted to requiring a signed attestation from freelancers plus a spot-check with Hive rather than scanning everything, which reduced the operational burden while maintaining accountability.

The detection tools are useful as a deterrent and a safety net, but building your entire trust framework around them is a mistake given their current limitations.

just want to echo that the process documentation approach works. I build WordPress and Shopify sites and occasionally get asked about AI usage for the custom graphics I create. Keeping organized Figma files with full history has been enough every time. No client has ever pushed back after seeing the actual design evolution.

Good thread. The data @voidvibes92 compiled saves a lot of individual research time. One thing I’d add from the agency management side is that most clients asking about AI detection right now aren’t doing it because they distrust their freelancers specifically. They’re responding to pressure from their own legal and compliance teams who read about AI art controversies and want a policy in place. Understanding that context helps when these conversations come up.

My approach has been to proactively include an “AI Tools and Methods” section in every project proposal, similar to what @CrispestHaze88 described in another thread. Listing exactly which tools I use, including AI-assisted ones, and how they fit into my process turns a potentially adversarial detection conversation into a transparent collaboration. Every client who’s received this has responded positively.