Frontier AI Still Does Not Read Sign Language. Data Is the Difference.

The best frontier model scored 0.112. Our model scored 0.779 on the same signs.

Frontier models write code, read X-rays and hold spoken conversations. So we asked a simple question: can they read American Sign Language?

We gave three of them our benchmark, Épée 200: GPT-6 Sol from OpenAI, and Claude Opus 5.5 and Claude Opus 5 from Anthropic. Then we compared them with Gallaudet 1.3, the recognition model we train on the CLERC corpus. It is the model behind our live demo.

The short answer: no, not yet. Our reading is that the gap is about data, not intelligence.

The test

Épée 200 is the open benchmark we published on clerc.io/benchmark. We used its Qualifier round:

  • 200 ASL sentences, 1,098 signs, 606 of them scored, over a vocabulary of 71 signs.
  • Five native Deaf signers.
  • The sign boundaries are given. The task is only to name each sign.
  • Every frontier model sees the same thing: 8 frames per sign, drawn as a stick figure from the public keypoints (body, hands, eyes, mouth and head outline). No video. Our signers' contracts do not allow their video to be sent to a third party, and the keypoints are what anyone can download.
  • Two questions per sign. First, pick the sign among the 71 options (multiple choice). Second, name it freely (open answer).
  • Scoring fixed in advance. Prompts, accepted answers and test items were hashed before the first call (protocol hash cdaf37c4).

Gallaudet 1.3 works from its own features and never gets the list of 71 options: it answers among its own 113 signs. The multiple choice therefore helps the frontier models, not ours. The open answer is the strict comparison.

The results

Macro-F1 on the 606 scored signs of the Épée 200 Qualifier: Gallaudet 1.3 0.779, linear baseline on public Épée data 0.179, GPT-6 Sol 0.112, Claude Opus 5.5 0.104, Claude Opus 5 0.047.

Multiple choice among the 71 signs, 606 scored signs. Guessing gives 1.4 %.

ModelMacro-F1Top-1Top-1, signer never seen by Gallaudet 1.3
Gallaudet 1.3 (CLERC)0.77978.0 %72.6 %
Linear baseline, public Épée data only0.17928.2 %37.6 %
GPT-6 Sol0.11212.5 %7.7 %
Claude Opus 5.50.10416.2 %12.0 %
Claude Opus 50.0479.6 %6.0 %

One of the five signers was never in Gallaudet 1.3's training data. On this signer, our model still reads 72.6 % of the signs. The frontier models read 6 to 12 %.

With a free answer, the frontier models name the right sign 4.0 % (GPT-6 Sol), 6.9 % (Claude Opus 5.5) and 4.8 % (Claude Opus 5) of the time.

Macro-F1 gives every sign the same weight, so a model cannot score well by only knowing the frequent ones. It is the ranking metric of the benchmark. The four models are live on the leaderboard; the linear baseline is the reference line under it.

What the models actually do

They do not read the signs. They fall back on a few default answers.

  • Claude Opus 5.5 answered MY 110 times. MY is the right answer 36 times.
  • GPT-6 Sol answered FINISH 104 times. FINISH is the right answer 13 times.
  • Claude Opus 5 answered MOTHER 89 times. MOTHER appears twice in the whole exam.

Between 43 % and 57 % of all their answers fall on just five signs. For comparison, answering I on every position, the most frequent sign of the exam, already gives 7.4 % top-1.

The hardest signs for them are the most common ones: pointing signs. I becomes MY or FEEL. YOU becomes MOTHER. The question sign becomes ONE.

Here is one of them, exactly as the frontier models saw it:

One sign from the Épée 200 Qualifier, 8 stick-figure frames. The true sign is YOU. Gallaudet 1.3 answers YOU. GPT-6 Sol, Claude Opus 5.5 and Claude Opus 5 all answer MOTHER.

Asked to name it freely, without the list, the three frontier models answered "mother", "eat" and "know".

Is the information missing from the input? No. The linear baseline sees the same keypoints, without even the depth, and gets it:

SignLinear baselineGPT-6 SolClaude Opus 5.5Claude Opus 5
YOU (39 times)28000
Question sign (30 times)29170

The signal is in the keypoints. The frontier models do not know how to read it.

Newer models progress, a little

Claude Opus 5.5 is the successor of Claude Opus 5. On the multiple choice, it doubles the macro-F1, from 0.047 to 0.104. The difference is beyond the benchmark's noise.

On the open answer, the two are not distinguishable (4.8 % against 6.9 %, overlapping intervals). And a good part of Opus 5.5's multiple-choice gain comes from its MY habit: leave out the positions where MY is the answer, and it reads 12.1 % of the signs, about the level of GPT-6 Sol (11.6 % without its own favourite).

A new model generation moved the needle by a few points. It did not change the picture.

Why: the data

The clearest line in the table is not the top one. It is the second.

A linear model, trained in seconds on the public Épée release alone (1,200 clips from 6 native Deaf signers), scores 0.179. That is more than every frontier model we tested. Gallaudet 1.3, trained on the CLERC corpus (more than ten hours of video, annotated sign by sign), scores 0.779.

We are not the first to see this. In April 2026, researchers at the University of West Bohemia tested GPT-5 and Gemini 2.5 Pro on WLASL300: 15.6 % and 23.7 %, far behind supervised models (Javorek et al., 2026).

Our reading: frontier models have read the internet, and the internet has very little annotated sign language, almost none of it from native Deaf signers with timed glosses. A small model with the right data beats a giant model without it.

Build on the data

That is the gap CLERC exists to close. We do not sell a model. We build the data layer underneath: native Deaf signers who signed an agreement and are paid for their work, recorded and annotated sign by sign by us, and delivered as keypoints with the annotations.

Gallaudet 1.3 is what that data does to a recognition model. If you build sign language AI, whether a frontier model, an avatar, a translation tool or a research corpus, you can license the same data. Explore the data or write to us.

Limits

  • Stick figures, not video. This is a handicap for models trained on photos and video. We chose it because our signers' contracts protect their image, and because it is drawn from the same keypoints every participant of Épée 200 gets.
  • Multiple choice helps the frontier models. They were given the 71 options; Gallaudet 1.3 was not.
  • Gallaudet 1.3 works from its own features, computed from the video, not from the stick figures.
  • Gallaudet 1.3 saw four of the five signers in training, never the exam clips. The fifth signer is the fair comparison. 43 of the 200 sentences were in its training pool, signed by someone else; this is stated on the benchmark page.
  • Models change fast. These results hold for the versions and dates below.

Try it

The 200 Qualifier clips are free to download as keypoints on clerc.io/benchmark, with a scorer that returns your card in seconds. If your model beats 0.179 on public data only, it tops the open data track.

Models tested: gpt-6-sol (Codex CLI), claude-opus-5-5 and claude-opus-5 (Claude Code CLI). Runs: September 28 and 29, 2026.

Reference. Javorek, Honzík, Gruber, Železný, Hrúz. "Sign Language Recognition in the Age of LLMs." arXiv:2604.11225, April 2026.

Follow @CLERC to track the build.

MORE FROM CLERC

September 23, 2026

The Names We Carry: Épée, Gallaudet, Clerc

Today is the International Day of Sign Languages. I dedicate it to the people who marked our history, because their names are written into every part of CLERC.

September 14, 2026

A Four-Second Avatar Is Not a Benchmark. That Is the Point.

SonZo AI just published the first external evaluation of CLERC data, under the 7 Production Gates. The avatar is the best result so far, and the review says it is a pilot, not a benchmark. Here is why that honesty is the whole collaboration.

August 18, 2026

Nobody Talks About the Language

Interpreters will not disappear, and fixing deafness is not the point. The real risk with sign language AI is quieter: a flattened, biased version of the language becoming the standard.

August 11, 2026

The Deaf Digital Heritage

For centuries, sign language lived only inside physical communities. The Deaf Digital Heritage is the structured, permanent corpus that changes that: built by the Deaf community, protected in Deaf hands, opened to the world.

August 3, 2026

Six Months That Beat Five Years

The last six months of CLERC produced more than the five years before them combined. The tools arrived for everyone at the same time. The difference is usage, and for Deaf founders the shift runs deeper than speed.

July 9, 2026

The Sign Language AI Dataset Landscape in 2026

WLASL, How2Sign, ASL Citizen, YouTube-SL-25, GoSign.AI: there have never been more sign language datasets. So why can't AI labs ship? An honest map of the landscape, and where the real gaps are.

July 6, 2026

Four Months, Less Than $1,000, and One Conviction: Data Comes First

CLERC's first public sign language recognition demo reached up to 71% accuracy on an unseen signer after four months, less than $1,000, and a Deaf-led data build.

June 22, 2026

Why Foundation Models Need ASL Training Data

Every major foundation model lab is missing the same thing: a sign language modality. Here is why ASL training data is not optional for the next generation of multimodal AI, and why generic video datasets do not close the gap.

June 1, 2026

Deaf People Use the Future First

Texting, captions, video calls. Deaf communities used them before the rest of the world caught up. Sign language is next, and this time people are not the only ones following.

April 27, 2026

Deaf Community First

Why CLERC exists because of, not for, the Deaf community and why that is a technical requirement, not a values statement.

April 22, 2026

We Are Not an Accessibility Company. We Are a Technology Company.

CLERC builds sign language AI infrastructure: structured, Deaf-led sign language data for ML labs, translation engines, and research teams. We're a tech company, not an accessibility company, and here's why that distinction matters.

April 13, 2026

From Raw Video to Labeled Signs

Raw footage in, production-ready sign language data out. How the CLERC pipeline transforms video into structured, linguistically grounded data.

April 7, 2026

Why Data, Not the Model

Sign Language AI does not have a model problem first. It has a data problem, and that changes everything.

April 1, 2026

THE CLERC DEMO: From Our Database to Your Use Case

Sign language has never had an ImageNet moment. CLERC is building the data foundation that makes sign language AI viable — and the demo shows exactly what that means in practice.

March 29, 2026

Manifesto of CLERC: The Foundation for the Next Generation of Sign Language AI

For nearly a decade, I have been obsessed with one challenge: breaking the communication barrier through technology.