INDUSTRIA
The VLM Detour - Why Juegos de CCTV Still Run on Purpose-Built Detectors in a Multimodal World
Roboflow refreshed its Vision Evals leaderboard on 10 July - the best vision-language model hits 61.7% mAP@50 on object detection. RF-DETR-2XL hits 78.5%. For a category built on precise real-time counting, that gap is why CCTV games run on a dedicated detector, not a large multimodal model.
Información fuente
Estado: Editorial
Fuente principal: equipo editorial de cctvgames.global
Last updated: 2026-07-30
Roboflow refreshed its Vision Evals leaderboard on 10 July, y el July 15 model recommendation is unambiguous - RF-DETR is still the strongest starting point for most computer vision projects in 2026, topping both COCO y el real-world RF100-VL benchmark. General-purpose vision-language models keep improving, but the gap to purpose-built detectors on object detection tasks remains 15 to 20 mAP points. For a category built on precise real-time counting, that gap is why CCTV games run on a dedicated detector and not on a large multimodal model.
The refreshed numbers
The Roboflow Vision Evals object-detection leaderboard, updated 10 July with pricing refreshed 16 July, ranks 16 vision-language models against 250 detection samples per model. The top result is Gemini 3.5 Flash at 61.7% mAP@50 and 37.9% mAP@75. GPT-5.6 Sol sits at 46.2 and 20.9. Claude Sonnet 5 at 18.0 and 4.4. These are strong absolute results for models that were not built for the task.
Compare to the Roboflow best-models roundup from 15 July. RF-DETR-2XL, released in 2026 as the first real-time detector to break 60 mAP@50:95 on COCO, scores 78.5% mAP@50 and 60.1% mAP@50:95 at 17 to 22 ms per frame. The medium variant lands 73.6% mAP@50 at 4.4 ms. On the stricter mAP@75 metric the gap is even wider - a purpose-built detector is roughly twice as accurate as the strongest general-purpose VLM.
Why the gap persists
Three structural reasons. Detección is fundamentally a geometric task - a model has to output tight bounding boxes with pixel-accurate coordinates. VLMs are trained to output tokens, y coordinate generation through tokens is inherently lossy. Detectors like RF-DETR are trained end-to-end to minimise localisation error directly.
Latency budgets differ by an order of magnitude. RF-DETR-2XL runs at 17.2 ms per frame on server GPU. Even the fastest VLM in the Vision Evals set - Gemini 3 Flash at 5.1 seconds per sample - is 300 times slower than a real-time detector. That is not a knob you tune. It is the difference between models that can process a video stream and models that can describe one still image.
Cost per inference is also off by orders of magnitude. Roboflow's July pricing for Claude Fable 5 lists USD 0.040 per detection sample. RF-DETR self-hosted on a T4 GPU processes 500 or more samples per second at fixed hardware cost. The Apostar Engine piece from June noted that operator-published games run on the operator's rails - and those rails cannot bear a per-frame VLM inference cost.
What this means for the count in Hora punta, río pato and carrera de nieve
Every round of Hora punta, río pato and carrera de nieve pushes a live surveillance feed through the same detection pipeline we covered in the RF-DETR versus YOLO26 piece from 6 July. The count that resolves the bet is a direct output of that detector's bounding boxes. Precision on that count is what the game sells - not narrative description, not scene explanation.
A VLM might one day narrate what happened in a round. It cannot decide what happened. For that, the model needs to output a stable, timestamped, framewise sequence of boxes that both the player y el regulator can inspect. The 2030 trillion GGR piece called this the explainable-AI angle - the detector's frame log is the fairness proof. A VLM's text description is not.
The CVPR benchmark that matters more
The 2026 CVPR paper ODOV introduced OD-LVIS, a 46,949-image benchmark spanning 15 real-world scenarios and 1,203 categories. This is a better proxy for CCTV feeds than COCO because it explicitly tests compound domain and category shifts. Watch for detector benchmarks on OD-LVIS through H2 2026 - that leaderboard, more than the VLM one, will tell us how CCTV pipelines age.
What to watch this quarter
Whether any VLM in the Vision Evals set crosses 70% mAP@50 by year-end. Whether RF-DETR-2XL adds an edge variant that trades accuracy for sub-10 ms latency on Jetson-class hardware. Whether OD-LVIS reveals a robustness gap that COCO-tuned models miss.
Where to play CCTV games remains Apostar, robet or Barajar. Responsible play matters more than the model architecture - set a session limit before you press start.
El juego implica riesgos: nunca apueste más de lo que puede permitirse perder.
Más noticias