Build an agent that answers questions about kitchen CCTV.
- 1BuildUse the starter brief.
- 2SubmitSend one version of your code.
- 3TestWe run the same published test.
- 4ScoreThe metric below sets your rank.
Entries are closed. Use the brief to study the task and results.Download starter brief ↓
About these results: These results cover six test questions. They do not establish performance on other kitchen footage.
The published September 7 deadline has passed. Ask about entry status before preparing a submission.
Review the sample videos and sample questions · download the plain-language starter brief · read the full challenge spec. Your program must return answers, supporting timestamps or frames, and a log of the run.
What to build
Build an agent we can run again. It receives kitchen videos and a JSON file of questions. It returns an answer for each question, the supporting timestamps or frames, and a run log. Nobody may inspect the videos manually during scoring.
python answer.py --videos ./videos --questions questions.json --out answers.json --log run_log.json
Find the moments that answer the question. Return not_visible when the footage does not show enough. The example command above is the preferred interface; the full specification explains the required files.
What to submit
Provide your repository or endpoint URL, agent name, models or APIs used, expected cost per scored run, and any setup notes. Your agent must produce answers, supporting timestamps or frames, and a run log.
The published deadline has passed. Check entry status with submit@builderr.ai first. The full specification explains the required files.
How we test it
We run your submitted agent on kitchen video with questions you have not seen before, then compare its answers with the answer key. It must find supporting moments in the footage without manual inspection during scoring.
For example, a question may ask when a handoff happened. Return the timestamp and supporting video evidence; if the footage does not show enough, return not_visible.
How you score
Total score is 100 points. A run is valid only if it stays under the hard cost and runtime caps. If it goes over, it does not rank.
- 80 points — correct answers: hidden yes/no, counts, event order, timestamp, duration, and not-visible questions.
- 15 points — lower model/API cost: lower model/API cost helps only after the answer is valid and useful.
- 5 points — a run we can repeat: one command, useful evidence, frame/model call log, and repeatable outputs.
Cost cap
Hard cap: $0.30 estimated model/API cost per 60 minutes of source video, 25 minutes wall-clock, and about 1,500 sampled frames per 60 minutes unless the evaluator states an equivalent frame budget.
If the scored set is longer or shorter, the dollar and frame caps scale with video minutes. The cap is deliberate: a useful restaurant tool has to find the right windows cheaply, not send every frame to the biggest model.
Build tips: finding the right moments
Start with a small end-to-end agent, then improve one question type at a time. Early reviewed entries suggest a question-driven search works better than a general video captioner. These are general starting guidelines, not a description of any entrant's implementation.
- Index the whole video cheaply, then inspect only likely windows.
- Route counts, states, OCR, timestamps, durations, and event order through separate paths.
- Crop the named person or object and use an upscaled region when the CCTV view is small.
- Confirm temporal hits on nearby frames and return
not_visiblewhen evidence is weak. - Log frames, model calls, runtime, cost, and evidence spans for every answer.
Sample videos and references
Start with the long fixed-view kitchen clips below. Scoring uses hidden questions and an answer key on the stated video set; if any held-back clips are added later, that will be stated before they are used.
| Long public CCTV | 60-minute fixed kitchen CCTV-style service clipOpen video -> | One-hour fixed kitchen view with timestamp, staff movement, prep, and service flow. Strong public sample/reference; verify rights before rehosting. |
| Long public CCTV | Cctv dapur keteter pesanan mulai jam 2 sampai jam 8 pagiOpen video -> | About 28 minutes of busy kitchen order flow from a fixed CCTV-style angle. Useful for hidden-question design and cost-control testing. |
| Long public CCTV | CCTV DAPUROpen video -> | About 22 minutes of fixed kitchen cooking footage. Good for long-window sampling and timestamp questions. |
| Long public CCTV | cctv dapur seafood jos gandosOpen video -> | About 19 minutes of restaurant-kitchen work from a fixed view. Useful as a public long-form reference. |
One-hour fixed kitchen view with timestamp, staff movement, prep, and service flow. Strong public sample/reference; verify rights before rehosting.
Open video ->About 28 minutes of busy kitchen order flow from a fixed CCTV-style angle. Useful for hidden-question design and cost-control testing.
Open video ->About 22 minutes of fixed kitchen cooking footage. Good for long-window sampling and timestamp questions.
Open video ->About 19 minutes of restaurant-kitchen work from a fixed view. Useful as a public long-form reference.
Open video ->More references and datasets
| Supplemental sample | Chinese Commercial Kitchen - overhead task clipsOpen source -> | Fixed-view real commercial-kitchen work. Useful supplemental sample and baseline development source. |
| Supplemental sample | Kaggle - Kitchen Video in RestaurantsOpen source -> | Restaurant-kitchen workflow footage. License/access must be checked before rehosting cuts. |
| Messy reference | Chinese restaurant hygiene-problem footageOpen video -> | Messier public restaurant footage. Useful for realism and question design, not the only scored source. |
| Style reference | IP Bullet CCTV Camera (Kitchen View) - Revlight SecurityOpen video -> | Short true fixed kitchen CCTV. Good visual target, too short to carry the round alone. |
| Reference only | COM KitchensOpen source -> | Unedited fixed-view cooking videos. Useful benchmark reference; restricted academic access. |
| Reference only | EPFL Smart KitchenOpen source -> | Long multi-view kitchen actions. Lab setting, but useful for scoring and action-question design. |
Fixed-view real commercial-kitchen work. Useful supplemental sample and baseline development source.
Open source ->Restaurant-kitchen workflow footage. License/access must be checked before rehosting cuts.
Open source ->Messier public restaurant footage. Useful for realism and question design, not the only scored source.
Open video ->Short true fixed kitchen CCTV. Good visual target, too short to carry the round alone.
Open video ->Unedited fixed-view cooking videos. Useful benchmark reference; restricted academic access.
Open source ->Long multi-view kitchen actions. Lab setting, but useful for scoring and action-question design.
Open source ->Sample questions
These are examples so builders understand the job. Final scoring uses hidden questions on the same public videos, with an answer key.
Model policy
Open-source or local models are recommended, because this should be cheap enough for a small kitchen to run often. They are not mandatory.
Cloud models are allowed if every call is logged and the normalized cost stays under the cap. Local runs count as $0 model/API cost because they use no paid model/API calls. This excludes local hardware and electricity costs. They still have to fit the same runtime and frame limits.
What the hidden questions test
Objective answers
Yes/no, multiple choice, counts, timestamps, durations, and not-visible answers. Freeform scene descriptions do not decide the winner.
Operational facts
Cap or hairnet visible? Container sealed before handoff? Tray unattended too long? Which station was active? When did a bottleneck start?
No guessing credit
If the order number, face, label, or action is not visible, the right answer is not visible. Guessing should lose points.
Temporal reasoning
Many questions need the order of events, not just a single frame: first sealed bag, last item added, duration unattended, or step before serving.
Benchmark to beat
The baseline is a simple coarse-to-fine video agent: sample the whole clip cheaply, identify likely time windows, inspect only those windows with a vision model and OCR where useful, then answer in JSON with evidence.
It must be cheap
A valid run stays under $0.30 estimated model/API cost per 60 minutes of source video. Full-video brute force may fit some cheap APIs, but repeated full-clip calls should not win.
It must be precise
The answer key checks timestamps, counts, event order, and not-visible cases. A broad caption or summary is not enough.
It must show evidence
Every useful answer should point to the video span or frame that supports it, with a run log showing frames, calls, cost, and runtime.
What a restaurant can learn from the result
A result shows how an entry answered the tested kitchen questions and what its model/API calls cost. The current six-question check is limited evidence; it does not establish accuracy across every kitchen, camera or operating condition. Use the linked videos and run evidence to understand that scope.