Frontier AI Still Does Not Read Sign Language. Data Is the Difference.
The best frontier model scored 0.112. Our model scored 0.779 on the same signs.
Frontier models write code, read X-rays and hold spoken conversations. So we asked a simple question: can they read American Sign Language?
We gave three of them our benchmark, Épée 200: GPT-6 Sol from OpenAI, and Claude Opus 5.5 and Claude Opus 5 from Anthropic. Then we compared them with Gallaudet 1.3, the recognition model we train on the CLERC corpus. It is the model behind our live demo.
The short answer: no, not yet. Our reading is that the gap is about data, not intelligence.
The test
Épée 200 is the open benchmark we published on clerc.io/benchmark. We used its Qualifier round:
- 200 ASL sentences, 1,098 signs, 606 of them scored, over a vocabulary of 71 signs.
- Five native Deaf signers.
- The sign boundaries are given. The task is only to name each sign.
- Every frontier model sees the same thing: 8 frames per sign, drawn as a stick figure from the public keypoints (body, hands, eyes, mouth and head outline). No video. Our signers' contracts do not allow their video to be sent to a third party, and the keypoints are what anyone can download.
- Two questions per sign. First, pick the sign among the 71 options (multiple choice). Second, name it freely (open answer).
- Scoring fixed in advance. Prompts, accepted answers and test items were hashed before the first call (protocol hash cdaf37c4).
Gallaudet 1.3 works from its own features and never gets the list of 71 options: it answers among its own 113 signs. The multiple choice therefore helps the frontier models, not ours. The open answer is the strict comparison.
The results

Multiple choice among the 71 signs, 606 scored signs. Guessing gives 1.4 %.
| Model | Macro-F1 | Top-1 | Top-1, signer never seen by Gallaudet 1.3 |
|---|---|---|---|
| Gallaudet 1.3 (CLERC) | 0.779 | 78.0 % | 72.6 % |
| Linear baseline, public Épée data only | 0.179 | 28.2 % | 37.6 % |
| GPT-6 Sol | 0.112 | 12.5 % | 7.7 % |
| Claude Opus 5.5 | 0.104 | 16.2 % | 12.0 % |
| Claude Opus 5 | 0.047 | 9.6 % | 6.0 % |
One of the five signers was never in Gallaudet 1.3's training data. On this signer, our model still reads 72.6 % of the signs. The frontier models read 6 to 12 %.
With a free answer, the frontier models name the right sign 4.0 % (GPT-6 Sol), 6.9 % (Claude Opus 5.5) and 4.8 % (Claude Opus 5) of the time.
Macro-F1 gives every sign the same weight, so a model cannot score well by only knowing the frequent ones. It is the ranking metric of the benchmark. The four models are live on the leaderboard; the linear baseline is the reference line under it.
What the models actually do
They do not read the signs. They fall back on a few default answers.
- Claude Opus 5.5 answered MY 110 times. MY is the right answer 36 times.
- GPT-6 Sol answered FINISH 104 times. FINISH is the right answer 13 times.
- Claude Opus 5 answered MOTHER 89 times. MOTHER appears twice in the whole exam.
Between 43 % and 57 % of all their answers fall on just five signs. For comparison, answering I on every position, the most frequent sign of the exam, already gives 7.4 % top-1.
The hardest signs for them are the most common ones: pointing signs. I becomes MY or FEEL. YOU becomes MOTHER. The question sign becomes ONE.
Here is one of them, exactly as the frontier models saw it:

Asked to name it freely, without the list, the three frontier models answered "mother", "eat" and "know".
Is the information missing from the input? No. The linear baseline sees the same keypoints, without even the depth, and gets it:
| Sign | Linear baseline | GPT-6 Sol | Claude Opus 5.5 | Claude Opus 5 |
|---|---|---|---|---|
| YOU (39 times) | 28 | 0 | 0 | 0 |
| Question sign (30 times) | 29 | 1 | 7 | 0 |
The signal is in the keypoints. The frontier models do not know how to read it.
Newer models progress, a little
Claude Opus 5.5 is the successor of Claude Opus 5. On the multiple choice, it doubles the macro-F1, from 0.047 to 0.104. The difference is beyond the benchmark's noise.
On the open answer, the two are not distinguishable (4.8 % against 6.9 %, overlapping intervals). And a good part of Opus 5.5's multiple-choice gain comes from its MY habit: leave out the positions where MY is the answer, and it reads 12.1 % of the signs, about the level of GPT-6 Sol (11.6 % without its own favourite).
A new model generation moved the needle by a few points. It did not change the picture.
Why: the data
The clearest line in the table is not the top one. It is the second.
A linear model, trained in seconds on the public Épée release alone (1,200 clips from 6 native Deaf signers), scores 0.179. That is more than every frontier model we tested. Gallaudet 1.3, trained on the CLERC corpus (more than ten hours of video, annotated sign by sign), scores 0.779.
We are not the first to see this. In April 2026, researchers at the University of West Bohemia tested GPT-5 and Gemini 2.5 Pro on WLASL300: 15.6 % and 23.7 %, far behind supervised models (Javorek et al., 2026).
Our reading: frontier models have read the internet, and the internet has very little annotated sign language, almost none of it from native Deaf signers with timed glosses. A small model with the right data beats a giant model without it.
Build on the data
That is the gap CLERC exists to close. We do not sell a model. We build the data layer underneath: native Deaf signers who signed an agreement and are paid for their work, recorded and annotated sign by sign by us, and delivered as keypoints with the annotations.
Gallaudet 1.3 is what that data does to a recognition model. If you build sign language AI, whether a frontier model, an avatar, a translation tool or a research corpus, you can license the same data. Explore the data or write to us.
Limits
- Stick figures, not video. This is a handicap for models trained on photos and video. We chose it because our signers' contracts protect their image, and because it is drawn from the same keypoints every participant of Épée 200 gets.
- Multiple choice helps the frontier models. They were given the 71 options; Gallaudet 1.3 was not.
- Gallaudet 1.3 works from its own features, computed from the video, not from the stick figures.
- Gallaudet 1.3 saw four of the five signers in training, never the exam clips. The fifth signer is the fair comparison. 43 of the 200 sentences were in its training pool, signed by someone else; this is stated on the benchmark page.
- Models change fast. These results hold for the versions and dates below.
Try it
The 200 Qualifier clips are free to download as keypoints on clerc.io/benchmark, with a scorer that returns your card in seconds. If your model beats 0.179 on public data only, it tops the open data track.
Models tested: gpt-6-sol (Codex CLI), claude-opus-5-5 and claude-opus-5 (Claude Code CLI). Runs: September 28 and 29, 2026.
Reference. Javorek, Honzík, Gruber, Železný, Hrúz. "Sign Language Recognition in the Age of LLMs." arXiv:2604.11225, April 2026.
Follow @CLERC to track the build.