The recordings might look sharp and the label files might open without errors. You still need to check whether the camera captured the action, whether the labels match what happened, and how much footage the supplier reviewed.

This guide is for engineering and data leads at robotics companies and AI labs buying human demonstration recordings, video annotation, or both. If you need robot control commands, sensor readings or calibration data, include those requirements separately in your brief.

The checks below help you agree what a supplier must deliver and decide whether to order more. The review sheet turns them into 15 checks with space for your requirements, findings and next steps.

Get the free pilot review sheet

1. The decision your pilot needs to support

Start with the task you need the data for. A pilot for recognising a hand movement may need different recordings and labels from one for judging whether a task succeeded.

Write down the required actions, settings and variations. Specify the amount of data, camera views, file formats and delivery date. Ask for examples of a usable recording and one your team would reject. This gives you something more concrete to discuss than a promise of “high-quality data.”

Agree the permitted uses of the recordings and who may access them, including any external AI tools. Record any limits on training, sharing or retention. Resolve missing instructions before collection or annotation begins.

Before work starts: agree the brief, acceptance rules and person who resolves questions. Record later changes so both teams review the same specification.

2. What the cameras actually show

Check whether the footage shows the evidence needed for each label. An overview may show the task, while a closer view reveals whether a hand released an object.

For recordings from several cameras, use an event visible in each view to check their timing. Equal recording lengths are not enough. Review any missing sections or gaps before judging actions near them.

Consider a fictional example: a person moves a cup onto a tray, but their hand hides the cup during the move. The file opens correctly. A label says the cup slipped at a precise time. That label still needs support from another aligned view, or it should remain uncertain.

Evidence to request: a sample covering the views your team needs, any camera timing differences, and a list of recordings with hidden or missing actions.

3. Labels and timestamps your team can check

A label needs a definition, an example and a rule for when it starts and ends. Include actions that look similar but should receive a different label. A deliberate release, for example, should not become an accidental drop just because the object leaves a hand.

If a label depends on repeated attempts, agree what counts as an attempt and when the sequence resets. For timing, agree how early or late a label may begin or end. Those limits should follow your task, not an unrelated project's rules.

Check a few labels against the original recordings. Each delivered label should point to the correct file, camera and moment. An edited or slowed review copy can have a different timeline. If the delivery uses frame numbers, confirm where numbering starts and whether the last frame belongs to the interval.

Evidence to request: dated label definitions, source recording IDs and a sample your engineer can trace back to the original footage.

4. Events the supplier may have missed

Reviewing the labels tells you whether the proposed events look right. It does not tell you which events the supplier failed to label.

Include sections with no labels in your review. Ask which parts of each recording received an initial check, which received closer inspection and which remain unchecked. To count missed events in a reference recording, a reviewer needs to check its complete timeline.

If AI helps produce the labels, ask how the system reads the video. A workflow that checks occasional frames can skip a brief event between them. Inspect disputed moments in the original footage, including the action before and after them. More sampled frames alone do not establish that a model interpreted the action correctly.

Evidence to request: a record of the footage reviewed and a check for missed events, including apparently uneventful sections.

5. Unclear labels and reviewer corrections

Ask the supplier to mark what the footage cannot establish. “The object is hidden behind the hand” supports a different conclusion from “the object was dropped.” Keep an unresolved label marked uncertain until the evidence supports a decision.

Keep the first annotation, the reviewer's correction and the final decision. The record should explain why a label changed and who resolved disagreements. If nobody has reviewed an item, its status should say so.

A buyer accepting a corrected delivery does not prove that the first set of AI-generated labels was accurate. Ask to see both versions if automation quality is part of the offer.

Evidence to request: unresolved cases and a correction record, with a named owner for each remaining question.

6. A quality check against independent references

Ask for reference answers that a reviewer has checked against the original footage. Use those answers to judge the supplier's labels. Comparing a tool's output with itself does not provide an independent quality check. Annotation tools such as CVAT support reference annotations, but the reference answers still need careful review. CVAT quality-control documentation

Keep examples used to improve instructions or models separate from the examples used for final evaluation. Using test information during development can make results look better than they are. scikit-learn guidance on data leakage

For a video pilot, we recommend grouping all views and edited copies of one recording together. Assign that whole group to either development or evaluation. Agree how to match a proposed event to a reference before counting errors, including the allowed timing difference and how to handle duplicate labels.

Ask for separate counts of missed events, false alarms (events that did not happen), wrong labels and timing errors. List unresolved cases too. Include the number of recordings and reference events reviewed. One accuracy percentage can conceal which problem your team would need to fix.

7. The work left after delivery

Import the sample into your team's tools. Check that the labels refer to the right recordings, timestamps fall inside those recordings and corrections remain in the exported files.

The delivery should include the file layout, label-guide version, review history and known limitations. Agree who fixes rejected items, whether corrections are included in the price, and when you will receive the revised files.

Before the pilot, agree how much checking and correction work your team can accept. Measure that effort on the sample and include it when comparing suppliers, alongside the quoted price. If the pilot is meant to improve a model, test that outcome separately. Accepting the data does not establish an improvement in model performance.

8. Your decision on the next batch

Use the agreed acceptance rules to decide whether to order more, request corrections or stop. For each unresolved issue, record the affected data, who will act and when you will review it again.

The review sheet gives you a place to record those decisions. Set your acceptance rules before review and use the same checks when comparing suppliers.

FREE BUYER WORKSHEET

Robotics data pilot review sheet

15 checks to use before ordering a larger batch. Record your acceptance rules, review the supplier’s evidence and decide what needs fixing.

Free, fillable PDF. Four pages.

We’ll use your details to handle this download request. No newsletter subscription. Privacy notice

A robotics data pilot with Legendre

Legendre offers human demonstration recordings and annotation or review of existing robotics video. Tell us what your team needs the data to support. We can discuss the recording scope, labels, review requirements and delivery format before agreeing a pilot.

Discuss your robotics data pilot