arrow_back EU collaborations & research

Open research proposal β€” seeking consortium partners

ETALON: an independent European
successor to Perspective API

On 31 December 2026, Google's Jigsaw shuts down Perspective API β€” the free toxicity classifier that over a thousand platforms, including major newsrooms, and a decade of hate-speech research have relied on. There is no migration path.

The obvious replacement β€” routing everything through closed LLM endpoints β€” repeats the same mistake: silent model updates, no model cards, no public documentation of what "toxic" even means.

In April 2026, researchers from the Weizenbaum Institute, TU Berlin, Oxford, KU Leuven and other European institutions published the ten requirements any legitimate successor must meet. ETALON is our proposal to help build it: an open, context-aware moderation model family β€” bringing a production system already running in Polish, a language almost no benchmark covers, plus the deployment layer academic groups don't have.

Status: early-stage proposal. We have not secured a coordinator or funding yet. If you're a research group, NGO, or institution working on this β€” we want to talk to you.

Dec 31, 2026

Perspective API shuts down

10

requirements the successor must meet

0

migration paths offered

Why this matters now

Three pressures are converging in the next eighteen months.

The incumbent instrument disappears.

Perspective API β€” free, used by over a thousand platforms including major newsrooms, covering 18 languages β€” shuts down on 31 December 2026 with no migration path. New quota requests have been closed since February 2026. An entire research literature loses its measurement instrument: Perspective-based benchmarks become impossible to verify or reproduce.

The obvious replacement repeats the mistake.

The default path β€” closed LLM endpoints as toxicity classifiers β€” reproduces the structural failures that made the shutdown so disruptive: silent model updates, construct underspecification, no model cards, no research-facing team. Outsourcing the measurement of toxicity to entities with active interests in its definition is not a neutral technical choice.

The regulatory demand is at its peak.

The Digital Services Act is fully applicable, with notice-and-action, statement-of-reasons and appeals duties that fall hardest on smaller platforms. The EU's Digital Omnibus on AI entered into force on 27 July 2026, setting documentation and transparency expectations for exactly this kind of system.

The specification

The ten requirements.

In April 2026, researchers from the Weizenbaum Institute, TU Berlin, Oxford, KU Leuven and other European institutions published "Bye Bye Perspective API," specifying what a legitimate successor must do. This is their specification, not our self-declaration β€” we're citing it, not authoring it.

T1

Reproducibility

Document all evaluation metadata β€” model, benchmark, metric β€” in a shared, machine-readable schema.

T2

Validity

Document explicitly which constructs, conditions and input formats are measured, including out-of-scope cases.

T3

Contextuality

Support structured context β€” sender, target, prior conversation β€” so counterspeech and reclaimed language can be told apart from abuse.

T4

Uncertainty

Return a full probability distribution over labels, not a single point estimate, since annotator disagreement reflects a genuinely contested construct.

T5

Multilinguality

Validate and publish performance per language, and disclose where it falls below a defined threshold.

G1

Independence

Open-source and institutionally independent, so no single organisation controls the schema, weights or definition of toxicity.

G2

Transparency

Open, auditable weights, training data and annotation guidelines, including annotator demographics.

G3

Legitimacy

Defining toxicity and selecting training data must involve affected communities through an open, documented process.

G4

Accountability

Continuous, public auditing with documented remediation across languages, demographics and content types.

G5

Sustainability

A funding model that does not reproduce the single-funder dependence that made Perspective's shutdown so disruptive.

Source: "Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation" (arXiv:2604.25580, April 2026)

What we propose to contribute

A production moderation system with real comments.

Academic groups have the mandate, the method and the credibility. What they don't have is a live moderation system, real deployment data, or Polish/CEE language coverage. This is what we'd bring to a consortium.

check_circle Context as a first-class input β€” parent post, thread history, community norms β€” not an ablation
check_circle A calibrated sensitivity threshold instead of an arbitrary cut-off
check_circle Polish and CEE-language coverage most benchmarks still ignore
check_circle A deployment layer: SDKs, a moderation dashboard, Digital Services Act transparency reports
check_circle A live pilot environment and real production data
check_circle An open-core commercial route that answers the sustainability requirement, not just gestures at it

The open research question

Custom-policy adherence β€” the sharpest gap, and ZenFeed's actual product problem.

An 8B open model already beats a frontier API on standard guardrail benchmarks β€” GuardReasoner-8B reaches 81.09% average F1 on prompt-harmfulness detection, ahead of GPT-4o+CoT. That part of the question is settled.

What isn't settled: on DynaBench, which tests whether a model follows a community's own policy rather than one fixed notion of harm, the same model scores 22.0 against 81.5 on standard safety benchmarks. Current guardrails encode one definition of harm and collapse when a community defines its own.

That is exactly ZenFeed's premise β€” a community administrator sets the threshold and the norms, not the model vendor β€” and exactly the capability the field currently lacks.

Research integrity

How we'd avoid repeating Perspective's mistake.

Perspective's scores were used to generate training labels, and models trained on those labels were then evaluated against Perspective's own outputs β€” so its systematic errors became embedded in the evaluation criterion and stayed invisible. Distilling from a frontier model and evaluating against benchmarks built on frontier-model labels would reproduce the same failure a generation later.

Design rule:

frontier models may bootstrap training data. They never define the construct, and they are never the evaluation ground truth. Agreement with a teacher model is reported as a diagnostic, never as a score.

Who we're looking for

Partners for a consortium that doesn't exist yet.

Research & institutional partners

Measurement infrastructure, annotation methodology, evaluation design.

Coordinator

A research group or institution to lead the consortium and its EU funding application.

Technology partners

Open European LLM development and infrastructure.

Funding routes we're evaluating

Horizon Europe Cluster 2 (democracy, measurement infrastructure) or Cluster 4 (open AI models); CERV-Daphne as a possible de-risking pilot; a research-infrastructure route linked to CLARIN. None of these is decided β€” we're still checking fit.

We're also scoping a complementary child-safety strand β€” privacy-preserving grooming-risk detection β€” as a separate proposal with its own consortium requirements.

Let's talk.

Download the Partner Identification Form, or write to us directly.