Results publishedkitchen CCTV

Build an agent that answers questions about kitchen CCTV.

For restaurants and cloud kitchens: answer whether staff wore caps, when a handoff happened, and what the camera could not show. Return evidence and stay under $0.30 per 60 minutes of video.
Prize
$300 to the winner
Completed
Sep 7
Scored on
Answer accuracy + cost
Model/API limit
$0.30 / 60 min cap
  1. 1BuildUse the starter brief.
  2. 2SubmitSend one version of your code.
  3. 3TestWe run the same published test.
  4. 4ScoreThe metric below sets your rank.

Entries are closed. Use the brief to study the task and results.Download starter brief ↓

About these results: These results cover six test questions. They do not establish performance on other kitchen footage.

The published September 7 deadline has passed. Ask about entry status before preparing a submission.

Review the sample videos and sample questions · download the plain-language starter brief · read the full challenge spec. Your program must return answers, supporting timestamps or frames, and a log of the run.

What to build

Build an agent we can run again. It receives kitchen videos and a JSON file of questions. It returns an answer for each question, the supporting timestamps or frames, and a run log. Nobody may inspect the videos manually during scoring.

python answer.py --videos ./videos --questions questions.json --out answers.json --log run_log.json

Find the moments that answer the question. Return not_visible when the footage does not show enough. The example command above is the preferred interface; the full specification explains the required files.

What to submit

Provide your repository or endpoint URL, agent name, models or APIs used, expected cost per scored run, and any setup notes. Your agent must produce answers, supporting timestamps or frames, and a run log.

The published deadline has passed. Check entry status with submit@builderr.ai first. The full specification explains the required files.

How we test it

We run your submitted agent on kitchen video with questions you have not seen before, then compare its answers with the answer key. It must find supporting moments in the footage without manual inspection during scoring.

For example, a question may ask when a handoff happened. Return the timestamp and supporting video evidence; if the footage does not show enough, return not_visible.

How you score

Total score is 100 points. A run is valid only if it stays under the hard cost and runtime caps. If it goes over, it does not rank.

  • 80 points — correct answers: hidden yes/no, counts, event order, timestamp, duration, and not-visible questions.
  • 15 points — lower model/API cost: lower model/API cost helps only after the answer is valid and useful.
  • 5 points — a run we can repeat: one command, useful evidence, frame/model call log, and repeatable outputs.

Cost cap

Hard cap: $0.30 estimated model/API cost per 60 minutes of source video, 25 minutes wall-clock, and about 1,500 sampled frames per 60 minutes unless the evaluator states an equivalent frame budget.

If the scored set is longer or shorter, the dollar and frame caps scale with video minutes. The cap is deliberate: a useful restaurant tool has to find the right windows cheaply, not send every frame to the biggest model.

Build tips: finding the right moments

Start with a small end-to-end agent, then improve one question type at a time. Early reviewed entries suggest a question-driven search works better than a general video captioner. These are general starting guidelines, not a description of any entrant's implementation.

  1. Index the whole video cheaply, then inspect only likely windows.
  2. Route counts, states, OCR, timestamps, durations, and event order through separate paths.
  3. Crop the named person or object and use an upscaled region when the CCTV view is small.
  4. Confirm temporal hits on nearby frames and return not_visible when evidence is weak.
  5. Log frames, model calls, runtime, cost, and evidence spans for every answer.

Sample videos and references

Start with the long fixed-view kitchen clips below. Scoring uses hidden questions and an answer key on the stated video set; if any held-back clips are added later, that will be stated before they are used.

Long public CCTV
60-minute fixed kitchen CCTV-style service clip

One-hour fixed kitchen view with timestamp, staff movement, prep, and service flow. Strong public sample/reference; verify rights before rehosting.

Open video ->
Long public CCTV
Cctv dapur keteter pesanan mulai jam 2 sampai jam 8 pagi

About 28 minutes of busy kitchen order flow from a fixed CCTV-style angle. Useful for hidden-question design and cost-control testing.

Open video ->
Long public CCTV
CCTV DAPUR

About 22 minutes of fixed kitchen cooking footage. Good for long-window sampling and timestamp questions.

Open video ->
Long public CCTV
cctv dapur seafood jos gandos

About 19 minutes of restaurant-kitchen work from a fixed view. Useful as a public long-form reference.

Open video ->
More references and datasets
Supplemental sample
Chinese Commercial Kitchen - overhead task clips

Fixed-view real commercial-kitchen work. Useful supplemental sample and baseline development source.

Open source ->
Supplemental sample
Kaggle - Kitchen Video in Restaurants

Restaurant-kitchen workflow footage. License/access must be checked before rehosting cuts.

Open source ->
Messy reference
Chinese restaurant hygiene-problem footage

Messier public restaurant footage. Useful for realism and question design, not the only scored source.

Open video ->
Style reference
IP Bullet CCTV Camera (Kitchen View) - Revlight Security

Short true fixed kitchen CCTV. Good visual target, too short to carry the round alone.

Open video ->
Reference only
COM Kitchens

Unedited fixed-view cooking videos. Useful benchmark reference; restricted academic access.

Open source ->
Reference only
EPFL Smart Kitchen

Long multi-view kitchen actions. Lab setting, but useful for scoring and action-question design.

Open source ->

Sample questions

These are examples so builders understand the job. Final scoring uses hidden questions on the same public videos, with an answer key.

Q1At what timestamp was the first sealed bag placed on the handoff shelf?
Q2Was the cook at the stove wearing a cap or hairnet at 00:45?
Q3How many people were active at the prep counter at 00:45?
Q4Did the worker close the container before moving it away from the station?
Q5Which happened last: garnish added, lid closed, bag moved, or tray wiped?
Q6Is the order number visible? Answer not_visible if it is not readable.

Model policy

Open-source or local models are recommended, because this should be cheap enough for a small kitchen to run often. They are not mandatory.

Cloud models are allowed if every call is logged and the normalized cost stays under the cap. Local runs count as $0 model/API cost because they use no paid model/API calls. This excludes local hardware and electricity costs. They still have to fit the same runtime and frame limits.

What the hidden questions test

Objective answers

Yes/no, multiple choice, counts, timestamps, durations, and not-visible answers. Freeform scene descriptions do not decide the winner.

Operational facts

Cap or hairnet visible? Container sealed before handoff? Tray unattended too long? Which station was active? When did a bottleneck start?

No guessing credit

If the order number, face, label, or action is not visible, the right answer is not visible. Guessing should lose points.

Temporal reasoning

Many questions need the order of events, not just a single frame: first sealed bag, last item added, duration unattended, or step before serving.

Benchmark to beat

The baseline is a simple coarse-to-fine video agent: sample the whole clip cheaply, identify likely time windows, inspect only those windows with a vision model and OCR where useful, then answer in JSON with evidence.

It must be cheap

A valid run stays under $0.30 estimated model/API cost per 60 minutes of source video. Full-video brute force may fit some cheap APIs, but repeated full-clip calls should not win.

It must be precise

The answer key checks timestamps, counts, event order, and not-visible cases. A broad caption or summary is not enough.

It must show evidence

Every useful answer should point to the video span or frame that supports it, with a run log showing frames, calls, cost, and runtime.

What a restaurant can learn from the result

A result shows how an entry answered the tested kitchen questions and what its model/API calls cost. The current six-question check is limited evidence; it does not establish accuracy across every kitchen, camera or operating condition. Use the linked videos and run evidence to understand that scope.

Ask about entry statusRead scoring rulesDownload challenge spec