Diese Seite enthält Affiliate-Links. Wir können von den Betreibern Provisionen erhalten, ohne dass für Sie Kosten entstehen. Erfahren Sie mehr

INDUSTRIE

The VLM Detour - Why CCTV Spiele Still Run on Purpose-Built Detectors in a Multimodal World

Roboflow refreshed its Vision Evals leaderboard on 10 July - the best vision-language model hits 61.7% mAP@50 on object detection. RF-DETR-2XL hits 78.5%. For a category built on precise real-time counting, that gap is why CCTV games run on a dedicated detector, not a large multimodal model.

Last updated: 30 July 2026
The VLM Detour - Why CCTV Spiele Still Run on Purpose-Built Detectors in a Multimodal World

Quellenangaben

Status: Redaktionell

Primärquelle: Redaktion von cctvgames.global

Last updated: 2026-07-30

Roboflow refreshed its Vision Evals leaderboard on 10 July, und die July 15 model recommendation is unambiguous - RF-DETR is still the strongest starting point for most computer vision projects in 2026, topping both COCO und die real-world RF100-VL benchmark. General-purpose vision-language models keep improving, but the gap to purpose-built detectors on object detection tasks remains 15 to 20 mAP points. For a category built on precise real-time counting, that gap is why CCTV games run on a dedicated detector and not on a large multimodal model.

The refreshed numbers

The Roboflow Vision Evals object-detection leaderboard, updated 10 July with pricing refreshed 16 July, ranks 16 vision-language models against 250 detection samples per model. The top result is Gemini 3.5 Flash at 61.7% mAP@50 and 37.9% mAP@75. GPT-5.6 Sol sits at 46.2 and 20.9. Claude Sonnet 5 at 18.0 and 4.4. These are strong absolute results for models that were not built for the task.

Compare to the Roboflow best-models roundup from 15 July. RF-DETR-2XL, released in 2026 as the first real-time detector to break 60 mAP@50:95 on COCO, scores 78.5% mAP@50 and 60.1% mAP@50:95 at 17 to 22 ms per frame. The medium variant lands 73.6% mAP@50 at 4.4 ms. On the stricter mAP@75 metric the gap is even wider - a purpose-built detector is roughly twice as accurate as the strongest general-purpose VLM.

Why the gap persists

Three structural reasons. Erkennung is fundamentally a geometric task - a model has to output tight bounding boxes with pixel-accurate coordinates. VLMs are trained to output tokens, Und coordinate generation through tokens is inherently lossy. Detectors like RF-DETR are trained end-to-end to minimise localisation error directly.

Latency budgets differ by an order of magnitude. RF-DETR-2XL runs at 17.2 ms per frame on server GPU. Even the fastest VLM in the Vision Evals set - Gemini 3 Flash at 5.1 seconds per sample - is 300 times slower than a real-time detector. That is not a knob you tune. It is the difference between models that can process a video stream and models that can describe one still image.

Cost per inference is also off by orders of magnitude. Roboflow's July pricing for Claude Fable 5 lists USD 0.040 per detection sample. RF-DETR self-hosted on a T4 GPU processes 500 or more samples per second at fixed hardware cost. The Einsatz Engine piece from June noted that operator-published games run on the operator's rails - and those rails cannot bear a per-frame VLM inference cost.

What this means for the count in Hauptverkehrszeit, Entenfluss and Schneelauf

Every round of Hauptverkehrszeit, Entenfluss and Schneelauf pushes a live surveillance feed through the same detection pipeline we covered in the RF-DETR versus YOLO26 piece from 6 July. The count that resolves the bet is a direct output of that detector's bounding boxes. Precision on that count is what the game sells - not narrative description, not scene explanation.

A VLM might one day narrate what happened in a round. It cannot decide what happened. For that, the model needs to output a stable, timestamped, framewise sequence of boxes that both the player und die regulator can inspect. The 2030 trillion GGR piece called this the explainable-AI angle - the detector's frame log is the fairness proof. A VLM's text description is not.

The CVPR benchmark that matters more

The 2026 CVPR paper ODOV introduced OD-LVIS, a 46,949-image benchmark spanning 15 real-world scenarios and 1,203 categories. This is a better proxy for CCTV feeds than COCO because it explicitly tests compound domain and category shifts. Watch for detector benchmarks on OD-LVIS through H2 2026 - that leaderboard, more than the VLM one, will tell us how CCTV pipelines age.

What to watch this quarter

Whether any VLM in the Vision Evals set crosses 70% mAP@50 by year-end. Whether RF-DETR-2XL adds an edge variant that trades accuracy for sub-10 ms latency on Jetson-class hardware. Whether OD-LVIS reveals a robustness gap that COCO-tuned models miss.

Where to play CCTV games remains Einsatz, Roobet or Shuffle. Responsible play matters more than the model architecture - set a session limit before you press start.

Glücksspiel ist mit Risiken verbunden – setzen Sie niemals mehr, als Sie sich leisten können, zu verlieren.

Weitere Neuigkeiten

Spielen Sie CCTV Spiele Where to Play

Bevor Sie gehen

Erhalten Sie einen Bonus auf Ihre erste Einzahlung beim führenden CCTV-Spieleanbieter.

NEUER BONUS
Claim Stake Bonus

18+ only. T&Cs apply. Gamble responsibly.