ÉPÉE 200 · 200 SENTENCES · 71 SIGNS SCORED · 2026 EDITION

An open benchmark for segmented ASL sign recognition.

Keypoints in, glosses out, every sign boundary supplied: it scores recognizing signs inside real sentences, not ASL recognition in general. Two rounds, clips from Deaf signers. The Qualifier scores your predictions in under a second: top-1, top-3, one held-out individual, BLEU. The Final is run by us, on sealed clips nobody has seen. That is the score we certify.

Evaluation only. By downloading you agree not to train on these clips and not to redistribute them. The 128 points, one by one.

clips: a Qualifier to try, a sealed Final to certify. None is in our open dataset
200 + 200clips: a Qualifier to try, a sealed Final to certify. None is in our open dataset
Deaf signers: four from the open dataset, on sentences they never released, plus one who is in no public release
5Deaf signers: four from the open dataset, on sentences they never released, plus one who is in no public release
signs scored, plus BLEU over the whole sentence
71signs scored, plus BLEU over the whole sentence
clips per round from one held-out individual (n=1), a person in no public release
44clips per round from one held-out individual (n=1), a person in no public release

01 / LEADERBOARD

Two rounds. Only the Final is certified.

200 clips that nobody outside CLERC has ever seen, in any form. You send the model, we run it, offline. Nobody can label what they cannot see: this is the only score we certify.

ÉPÉE 200 · FINAL · SEGMENTED RECOGNITION · BOUNDARIES SUPPLIED · LABELS-V1 · SINGLE DEAF ANNOTATOR

#MODELMACRO-F1TOP-1TOP-3HELD-OUT INDIVIDUAL (n=1)BLEU-1 · GERRUN
01

Gallaudet 1.3

PRIVATE DATACLERC · CLERC corpus (private, 10+ h of video)

Run by the organizer, internal features

Details
  • Macro-F1 on the signs seen 3+ times: 0.769, not ranked
  • False alarms on signs outside the vocabulary: 62.2%
  • Abstention, not ranked: 85.8% right on its 80% most confident signs, 90% accuracy up to 68.0% of the signs, AURC 0.064, false alarms 37.2% at 80% coverage
  • scorer 2.1 · epee200_v2_final · labels-v1 · single Deaf annotator · clips sha256 07382a05bcfd
0.775

qualifier 0.779 · consistent

77.7%87.4%76.8%†±6.943.8GER 56.4%ORGANIZER

Every row carries its training data: OPEN DATA = the public Épée release only, PRIVATE DATA = private or licensed CLERC data, OTHER DATA = anything else, as declared by the entrant. Reference, not ranked: baseline.py, trained on the open data only, scores 0.198. Ranked by macro-F1 over 71 signs. Rows closer than 3.8 points are a tie. Held-out individual (n=1): top-1 on one person who is absent from the open dataset, 44 clips, with its 95% interval (about ±7 points). It tells you about that one person, not about unseen signers in general. † that row trained on data that include this person, so for it the number is not an unseen-signer score. BLEU-1 and GER (gloss error rate, lower is better): whole sentence, open vocabulary. Details: macro-F1 on the signs seen 3+ times, false alarms, abstention, and the scorer, exam and clips each number was computed with. n/a: not measured on that run. Both rounds share signers, sentence lengths and vocabulary: a Qualifier-to-Final gap above 5.4 points is flagged for review.

What our own row has seen: Gallaudet 1.3 never trained on an exam clip, but 43 of the 200 Qualifier sentences (42 in the Final) were in its training pool, signed by someone else. Its top-1 is 86% on those sentences and 76% on new ones. No exam sentence is in the open dataset.

02 / ENTER

Qualify in five minutes. Certify in the Final.

ROUND 1 · QUALIFIER · INSTANT, SELF-REPORTED

  1. 1 · DOWNLOAD

    The 200 Qualifier clips in one zip (keypoints, 36 MB, sign boundaries given, no labels). Train on Épée v0.3 or on your own data.

  2. 2 · RUN

    One gloss per sign, or three ranked candidates if you want a top-3. Add a confidence per sign, or answer UNKNOWN, to be scored on abstention too. The baseline below does the whole loop: a few minutes of download the first time, then seconds.

  3. 3 · UPLOAD

    Drop predictions.json. Full score card in under a second, no account. To be listed, give a model name and an email: CLERC reviews every row first.

pip install numpy scikit-learn huggingface_hub
curl -O https://clerc.io/benchmark/baseline.py
python baseline.py        # fetches the data, trains on Épée v0.3, predicts the 200 clips

Drop predictions.json here

or click to browse · JSON, 200 clips

Your file is scored and discarded, only the scores and the names you type are kept. Prefer email? Send predictions.json to florian@clerc.io and we reply with the same card.

ROUND 2 · FINAL · CERTIFIED

You send the model, not the results.

  1. 1. Package your model with the submission kit (a Docker image, about an hour).
  2. 2. CLERC runs it on 200 sealed clips: same format as the Qualifier, no network, read-only, a few minutes.
  3. 3. Your row is published as Final · certified, next to your Qualifier score. If the two disagree beyond noise, the table shows it.
Request the kit: florian@clerc.io

03 / FINE PRINT

How it is scored

Same scoring in both rounds. Predictions align by position with the reference signs of each clip. A missing position counts as wrong. Only complete runs of 200 clips are scored, and the labels never leave our server.

Macro-F1 (the ranking): F1 per sign, averaged over the 71 scored signs, so a frequent sign like YOU cannot carry a model that fails on the rare ones. Support is uneven, from 1 to 45 tokens per sign in the Qualifier and 6 at the median, which is why the noise floor below is wide. Top-1 and top-3 are computed on the same 606 sign tokens. The score card also gives top-1 by sentence length.

Next to it, the score card and each row's details give macro-F1 over the signs seen 3 times or more (59 of the 71 in the Qualifier, 59 of the 70 in the Final), where one token can no longer swing a sign from 0 to 1. We measured whether it is steadier, and it is not: on our own model its bootstrap spread is 0.020 against 0.019, and its Qualifier-to-Final gap 0.030 against 0.004, because fewer signs are averaged. So the ranking stays on every sign, and the thresholded score is there to read, not to rank.

Held-out individual (n=1): top-1 on the one person who is absent from the public Épée v0.3 release, 44 clips and 154 scored signs in the Qualifier (142 in the Final). One person is not a population: it is evidence about this individual, within about ±7 points, not performance on unseen signers in general. And it is an unseen-signer score only for a model that never trained on private CLERC data. That person is in our private corpus and in the data we license, so rows trained on either, ours included, carry a † on that number. The next edition is planned with more held-out people. Épée 200 is an external, sealed exam: it complements a signer-disjoint evaluation on your own data, such as leave-one-signer-out, and does not replace it.

BLEU-1 and BLEU-2 compare your whole gloss sequence with the reference, every position included, open vocabulary: they reward knowing more than 71 signs. Sentences are too short for a stable BLEU-4.

Gloss error rate (GER, what the field calls WER): edit distance between your gloss sequence and the reference, over all 1,098 positions of the Qualifier, open vocabulary, lower is better. With boundaries given there are no insertions or deletions, so it is one minus the accuracy over every position, the 492 outside the scored vocabulary included. Our own model is at 54.6%: most of what it misses are signs it never learned, fingerspelled words and names.

False alarms: 492 positions of the Qualifier hold a sign outside the 71 (529 in the Final). They never count for top-1, but answering one of the 71 there is a false alarm, and the score card reports their share, lower is better. A closed vocabulary cannot avoid them: the public baseline, which knows only the 71, is at 100%, and our own model answers one of the 71 on 59% of those signs.

Abstention, reported and not ranked: give a confidence between 0 and 1 for each sign, or answer UNKNOWN. The card then orders the scored signs from most to least confident and reports the accuracy on the 80% most confident, the largest share answered at 90% accuracy, the area under the risk-coverage curve (AURC), and the false alarms left at 80% coverage. UNKNOWN counts as the least confident answer and as wrong. Equal confidences share their accuracy, so the order of your file never matters and a flat confidence earns nothing.

Noise floor: ±3.8 points of macro-F1 (bootstrap, 1,000 clip-level resamples). Rows inside it are marked as ties.

Every score card names the scorer version (2.1), the exam version, the label version (labels-v1: a single Deaf annotator) and the sha256 of its clips. A certified row adds the digest of the Docker image that ran and the Docker version, so each row can be traced to exactly what produced it. A new label version never rescores history in silence: existing rows keep the versions they were scored on, and results on new labels enter as new rows.

What it does not measure

Sign boundaries are given: this is recognition of segmented signs inside sentences. It is not segmentation, not continuous recognition and not translation, and a score here says nothing about ASL grammar: non-manual markers, use of space and classifiers are not evaluated, and fingerspelled words (52 of the 1,098 positions in the Qualifier) count toward BLEU only.

Five signers, elicited sentences of 5.5 signs on average (3.2 in the open dataset), frontal webcam framing: it is not an in-the-wild test, and it is not evidence that a system is ready for Deaf users. The public track is keypoints only: 128 points per frame, no pixels, 28 face points and a head outline instead of the face mesh.

The ranking is forced choice: every position must be answered, and in the rank a confident error costs the same as a guess. Abstention is measured beside it (see How it is scored) and stays out of the rank for this edition, so that every row is ranked on the same rule.

Why only the Final is certified

Hidden labels do not make a score trustworthy: whoever holds the clips can label them by hand in an afternoon, or hide those answers inside a container. Hidden inputs do. The clips of the Final have never left CLERC in any form: no video, no keypoints, and no exam clip of either round was ever part of a dataset we delivered to a customer.

The Final mirrors the Qualifier: same five signers, same number of clips per signer, same sentence lengths, same scored vocabulary (the Final carries 70 of the 71 signs), no clip and no sentence in common. Our own models score the same on both (0.779 and 0.775 for Gallaudet 1.3, 0.179 and 0.198 for the public baseline). A gap above 5.4 points is more than the clip draw explains, and the table flags it for review. A flag is not a verdict: overfitting to the public clips, sensitivity to the draw, calibration or an implementation difference can all open a gap.

CLERC is both organizer and entrant. Our own rows are therefore marked organizer, never certified, and the sealed set is identified by a published hash: sha256 07382a05bcfd. Our model family is named Gallaudet as a tribute: CLERC is not affiliated with Gallaudet University.

What our own row has seen: Gallaudet 1.3 never trained on an exam clip, but 43 of the 200 Qualifier sentences (42 in the Final) were in its private training pool, signed by another signer. Its top-1 is 86% on those and 76% on the new ones. No exam sentence is in the public release, so nobody on the open data track has that help. Customers who train on licensed CLERC data are in the same position as our own model, and their rows say so.

One of the four other signers was never in the training pool of Gallaudet 1.3. Its top-1 on that signer is 73% in the Qualifier, against 76% to 88% on the signers it trained on: that is its real unseen-signer number.

The sealed set is dedicated to Jean Massieu, Deaf teacher and the first teacher of Laurent Clerc. In London in 1815, Massieu and Clerc answered in public any question the audience put to them. Thomas Gallaudet was in that audience.

Sign by sign: our own model

The public score card gives no per-sign result and no confusion matrix: every extra number would help read the labels out of the score. We publish ours instead, and send an entrant theirs privately after a Final run.

Gallaudet 1.3 in the Qualifier, F1 on the 29 signs that occur at least 8 times. Best: WORK 1.00, NOW 0.95, CAN 0.95, OLD 0.94, MY 0.92, IN 0.90. Worst: WE 0.00, OR 0.00, SHE 0.27, POINTER_RIGHT 0.48, NEED 0.56, THIS 0.67.

Why: three scored signs (HE, OR, WE) are outside its 113-sign vocabulary, so it scores zero on them by construction, about 3.4 points of macro-F1. It answers I for WE 11 times out of 14. HE and SHE are the same pointing sign: gender is in the context, not in the keypoints. POINTER_RIGHT, another pointing sign, is read as NEED, SHE or 2. Of its 133 errors, 105 land on another scored sign and 28 outside the 71.

Abstention, same model: answering only its 80% most confident signs, Gallaudet 1.3 is right on 85.8%, against 78.0% when forced, and it holds 90% accuracy on up to 65% of the signs (68% in the Final). Signs outside the 71 are its weak spot: forced, it answers one of the 71 on 59% of them, and still on 36% at 80% coverage. Its confidence separates its right answers from its wrong ones (AUROC 0.81) better than from the signs outside the 71 (0.73).

Data statement

Who signs: five Deaf signers. Each signed a recording agreement that covers this use, and consent is recorded per clip. Faces are never published and signers are never named: the release is keypoints only.

How it was recorded: webcam, 1280 x 720, 30 frames per second, one signer facing the camera. Each clip renders one written English sentence in ASL: the prompts are sentences to translate, not scenarios or pictures to describe, and this is not spontaneous conversation. The recording screen shows the English sentence only, with no written instruction to sign word for word or freely.

The labels are what was signed, not the English: about 0.85 signs per English word, and about a third of the signs have no counterpart in the English sentence, even allowing for word endings (pointing, fingerspelling, numbers, and signs whose gloss is not the English word).

Labels: glosses and sign boundaries were annotated in-house by one Deaf annotator, so every score here measures agreement with one person's labels. Inter-annotator agreement has not been measured yet. Next: a second Deaf annotator relabels the 400 exam clips independently, blind to the first labels and to any model suggestion. We will publish agreement on glosses and on boundaries, adjudicate every disagreement, and release the adjudicated labels as labels-v2. Today's references are labels-v1, and every row and score card says so.

What the 128 points leave out: eyebrows, cheeks, eye gaze and most of the face mesh. Most non-manual grammar, such as question and negation marking or adverbial mouth shapes, is therefore not in the input, only eye opening, lips and head position are. Hand detail is missing whenever the tracker loses a hand: both hands are present on 36% of the frames only.

What five signers cannot represent: regional, generational, racial and ethnic variation in ASL. We do not publish signer demographics, and no claim on this page extends beyond these five people.

Intended use: comparing sign classifiers on identical inputs, with a number somebody else has checked. Not intended: claiming that a product understands ASL, or replacing evaluation with Deaf users.

Rules

The Qualifier clips are evaluation-only: no training, no fine-tuning, no augmentation on them, no redistribution. 5 scored runs per hour, 20 per day. Qualifier rows are reviewed by CLERC before they appear, and stay marked self-reported. A Qualifier row that would take first place is listed once a Final run confirms it, and CLERC may ask any entrant for a Final run before listing a row.

A Final · certified run is a Docker image executed by CLERC with no network and a read-only filesystem, on keypoints in the exact format of the Qualifier. We keep the predictions of every run, never the image. Three certified runs per team per edition: the Final is an exam, not a development set.

Open data track: models trained on the public Épée release only, the same 1,200 clips for everyone. Training data is declared by the entrant, and CLERC may ask for the training recipe before a row is listed there.

More signers in, better scores out.

The open dataset is 1,200 clips. The top row trained on more than ten hours of private video. If you need that to train, see the data or write to florian@clerc.io.