Type a slur into a comment box under a post, and a good content filter will catch it before anyone reads it. Paste the exact same slur into a screenshot or a meme caption instead, and some filters will wave it right through. Not because the content became less toxic. Because most moderation systems have spent years reading text — and to them, an image is just a file.
This isn’t a theoretical curiosity. It’s a documented, researched gap — and one of the reasons ZenFeed now transcribes and moderates image comments exactly the way it moderates text.
Why it works: a text filter can’t see text inside an image
A team at the Chinese University of Hong Kong tested this experimentally. They built a framework called OASIS which, using 21 transform rules derived from an analysis of 5,000 real toxic posts from Twitter, Instagram, Sina Weibo and Baidu Tieba, turns toxic text into an image while preserving its full meaning (for example, as a screenshot with light distortion). They ran the resulting images through five commercial content filters: Google Cloud, Microsoft Azure, Baidu Cloud, Alibaba Cloud and Tencent Cloud. The result: up to 100% of the generated images bypassed the filter, despite carrying the exact same toxic content that the very same systems flagged flawlessly as plain text.
This isn’t one vendor’s bug — it’s a structural gap. Analyzing a comment’s text and analyzing an image attached to that comment have historically been two entirely separate systems, run separately, if at all. If an abusive word lives in pixels rather than in a message field, a classic content filter simply doesn’t see it.
It’s not just slurs — memes carry disinformation too
The same gap works in the other direction as well: it doesn’t just hide profanity, it also lets memes and images carry false claims past the scrutiny that text is normally subject to. A systematic review published in December 2025 in JMIR Infodemiology (386 records screened, 14 peer-reviewed studies from 2020–2025) describes memes as “emotionally salient and visually potent” carriers of health narratives — effective precisely because humor and disgust bypass critical thinking faster than a plainly worded sentence.
A study by Dr Stephanie A. Baker and Dr Michael James Walsh of City St George’s, University of London (February 2024) shows this with a concrete example: popular anti-vaccine influencers during the COVID-19 pandemic systematically used memes to amplify and monetize the anti-vaccine movement, citing a finding that just a dozen or so (12) creators were responsible for a significant share of that content across social media.
The same tool doubles as a political weapon — and an organized one. In its report “Visual assessment of CIB in disinformation campaigns,” EU DisinfoLab applied 50 coordinated-inauthentic-behavior indicators to three campaigns built largely around imagery: “Operation Overload,” a Russian influence operation on TikTok (with DFRLab and BBC Verify), and QAnon’s “Save the Children” campaign. This isn’t lone bad actors — these are organized operations for which an image is simply a more effective vehicle than a sentence, because it’s harder to verify and easier to flood a channel with.
It’s hard for AI too — not just for a banned-word list
It’s worth being honest here: automated image analysis on its own doesn’t solve the problem overnight either. Meta AI demonstrated this back in 2020 with the Hateful Memes Challenge — a set of over 10,000 memes labeled hateful or not, with a $100,000 prize for the best model. The dataset deliberately includes pairs where neither the text nor the image alone is offensive — only together. The textbook example from the research paper: the caption “love the way you smell today” is neutral, a photo of a skunk is neutral, but combined they become an offensive joke. Models that judge only text or only the image have no way to catch this — they have to reason across both modalities at once. In the original benchmark, the best contemporary model (Visual BERT COCO) reached about 69.5% accuracy against 84.7% for trained human annotators. This remains an active research area, not a solved problem — just as context decides how plain text gets judged, context also decides how an image gets judged.
What we’re changing in ZenFeed
Until now, an image comment — a photo, a meme, or a screenshot sent without accompanying text, or alongside it — passed through ZenFeed with no analysis at all: since there was no text content to assess, the system simply skipped it. That’s exactly the gap the OASIS research above describes, just on our side of the pipeline instead of a bad actor’s.
Now ZenFeed’s worker transcribes the image attached to a comment (a verbatim reading in the original language, plus a short English description of the scene) and runs that output through the exact same toxicity assessment that scores plain text — including any automatic hide or delete action. In the Moderation panel, the transcript and description are visible under an “Image comment” heading, right next to the score.
A few specifics: the feature is available on the Starter and Pro plans (not on the Free plan), and it costs a flat extra 1 credit per image analyzed — regardless of whether the action is hide or delete, and regardless of whether ZenFeed’s auto-comment is enabled. Transcripts and descriptions follow the same retention rules as comment text: they disappear after 30 days along with the author’s data.
This isn’t a silver bullet — a heavily distorted image, or irony embedded in the picture itself, is still at the edge of the model’s accuracy, just as sarcasm is at the edge of accuracy for plain text. But it closes a gap that is today documented, measured, and — as the OASIS research shows — exploited with a success rate reaching 100%.
Want image analysis on your page?
Image-comment analysis runs automatically on Starter and Pro accounts — there’s nothing to configure. Pricing and limit details are on the pricing page. Registered NGOs and independent newsrooms get ZenFeed’s full feature set, including image analysis, for free through our non-profit program. Questions? Write to us at hello@zenfeed.eu.
Key takeaways
- Embedding toxic text into an image instead of typing it into a comment box is a documented way to evade filters — the OASIS study (ASE 2023) used it to bypass five commercial content filters (Google, Microsoft, Baidu, Alibaba, Tencent) with up to 100% success.
- The same gap doubles as a disinformation vector: a 2025 JMIR Infodemiology review describes memes as an effective carrier of health narratives, and a 2024 City St George’s study shows this in the anti-vaccine movement.
- Image-based disinformation can also be organized — EU DisinfoLab documented this across three concrete coordinated-inauthentic-behavior campaigns.
- Automated image analysis isn’t trivial either — Meta’s Hateful Memes Challenge (2020) showed the best contemporary model still clearly trailed humans (69.5% vs. 84.7%).
- ZenFeed now transcribes image comments and moderates them with the same toxicity assessment as text; the feature is available on Starter and Pro for a flat extra 1 credit per comment.
Sources
- Wang et al. – An Image is Worth a Thousand Toxic Words: A Metamorphic Testing Framework for Content Moderation Software (ASE 2023)
- JMIR Infodemiology – Internet Memes as Drivers of Health Narratives and Infodemics: Integrative Review (December 2025)
- City St George’s, University of London – How memes transformed from pics of cute cats to health disinformation super-spreaders (February 2024)
- EU DisinfoLab – Visual assessment of CIB in disinformation campaigns
- Meta AI – Hateful Memes Challenge and dataset
- Kiela et al. – The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes (NeurIPS 2020)