อุตสาหกรรม
The VLM Detour - Why เกมส์กล้องวงจรปิด Still Run on Purpose-Built Detectors in a Multimodal World
Roboflow refreshed its Vision Evals leaderboard on 10 July - the best vision-language model hits 61.7% mAP@50 on object detection. RF-DETR-2XL hits 78.5%. For a category built on precise real-time counting, that gap is why CCTV games run on a dedicated detector, not a large multimodal model.
แหล่งข้อมูล
สถานะ: บทบรรณาธิการ
แหล่งที่มาหลัก: ทีมบรรณาธิการ cctvgames.global
Last updated: 2026-07-30
Roboflow refreshed its Vision Evals leaderboard on 10 July, และ July 15 model recommendation is unambiguous - RF-DETR is still the strongest starting point for most computer vision projects in 2026, topping both COCO และ real-world RF100-VL benchmark. General-purpose vision-language models keep improving, but the gap to purpose-built detectors on object detection tasks remains 15 to 20 mAP points. For a category built on precise real-time counting, that gap is why CCTV games run on a dedicated detector and not on a large multimodal model.
The refreshed numbers
The Roboflow Vision Evals object-detection leaderboard, updated 10 July with pricing refreshed 16 July, ranks 16 vision-language models against 250 detection samples per model. The top result is Gemini 3.5 Flash at 61.7% mAP@50 and 37.9% mAP@75. GPT-5.6 Sol sits at 46.2 and 20.9. Claude Sonnet 5 at 18.0 and 4.4. These are strong absolute results for models that were not built for the task.
Compare to the Roboflow best-models roundup from 15 July. RF-DETR-2XL, released in 2026 as the first real-time detector to break 60 mAP@50:95 on COCO, scores 78.5% mAP@50 and 60.1% mAP@50:95 at 17 to 22 ms per frame. The medium variant lands 73.6% mAP@50 at 4.4 ms. On the stricter mAP@75 metric the gap is even wider - a purpose-built detector is roughly twice as accurate as the strongest general-purpose VLM.
Why the gap persists
Three structural reasons. การตรวจจับ is fundamentally a geometric task - a model has to output tight bounding boxes with pixel-accurate coordinates. VLMs are trained to output tokens, และ coordinate generation through tokens is inherently lossy. Detectors like RF-DETR are trained end-to-end to minimise localisation error directly.
Latency budgets differ by an order of magnitude. RF-DETR-2XL runs at 17.2 ms per frame on server GPU. Even the fastest VLM in the Vision Evals set - Gemini 3 Flash at 5.1 seconds per sample - is 300 times slower than a real-time detector. That is not a knob you tune. It is the difference between models that can process a video stream and models that can describe one still image.
Cost per inference is also off by orders of magnitude. Roboflow's July pricing for Claude Fable 5 lists USD 0.040 per detection sample. RF-DETR self-hosted on a T4 GPU processes 500 or more samples per second at fixed hardware cost. The เดิมพัน Engine piece from June noted that operator-published games run on the operator's rails - and those rails cannot bear a per-frame VLM inference cost.
What this means for the count in ชั่วโมงเร่งด่วน, แม่น้ำเป็ด and สโนว์รัน
Every round of ชั่วโมงเร่งด่วน, แม่น้ำเป็ด and สโนว์รัน pushes a live surveillance feed through the same detection pipeline we covered in the RF-DETR versus YOLO26 piece from 6 July. The count that resolves the bet is a direct output of that detector's bounding boxes. Precision on that count is what the game sells - not narrative description, not scene explanation.
A VLM might one day narrate what happened in a round. It cannot decide what happened. For that, the model needs to output a stable, timestamped, framewise sequence of boxes that both the player และ regulator can inspect. The 2030 trillion GGR piece called this the explainable-AI angle - the detector's frame log is the fairness proof. A VLM's text description is not.
The CVPR benchmark that matters more
The 2026 CVPR paper ODOV introduced OD-LVIS, a 46,949-image benchmark spanning 15 real-world scenarios and 1,203 categories. This is a better proxy for CCTV feeds than COCO because it explicitly tests compound domain and category shifts. Watch for detector benchmarks on OD-LVIS through H2 2026 - that leaderboard, more than the VLM one, will tell us how CCTV pipelines age.
What to watch this quarter
Whether any VLM in the Vision Evals set crosses 70% mAP@50 by year-end. Whether RF-DETR-2XL adds an edge variant that trades accuracy for sub-10 ms latency on Jetson-class hardware. Whether OD-LVIS reveals a robustness gap that COCO-tuned models miss.
Where to play CCTV games remains เดิมพัน, รูเบท or สับเปลี่ยน. Responsible play matters more than the model architecture - set a session limit before you press start.
การพนันมีความเสี่ยง - อย่าเดิมพันเกินกว่าที่คุณจะสูญเสียได้
ข่าวเพิ่มเติม