Why surveillance teams are moving to AI video analytics
A typical control room asks one operator to watch a wall of twenty feeds. Research on vigilance is unkind here: detection rates for small, brief events fall sharply after the first twenty minutes. The result is a paradox — more cameras, less certainty. AI video analytics changes the economics by splitting the job. The machine watches everything, continuously, and surfaces the small number of moments that deserve human attention. The operator stops being a screen-watcher and becomes a decision-maker.
The practical benefits show up in three places:
- Triage. Alerts arrive pre-filtered by type, location, and confidence, so the queue is ranked instead of random.
- Search. Footage becomes a queryable database. Instead of scrubbing timelines, you ask for "person in a hi-vis vest crossing the perimeter line between 01:00 and 03:00."
- Patterns. Aggregated event data exposes recurring problems: a gate propped open every Friday, a loading bay that attracts loitering after closing, a door that never quite latches.
None of this requires replacing your cameras overnight. Most successful rollouts start with analytics on a subset of existing streams, measure what the models actually catch, and expand only where the signal is clean. The rest of this guide walks through that process end to end — from architecture and pipeline design to cost control, privacy, and the mistakes that quietly kill projects.
The core building blocks of a modern analytics stack
Before evaluating a single model, it helps to map the stack into four layers. Problems that look like "the AI is bad" are often transport, storage, or camera problems wearing an AI costume.
Cameras, lenses, and edge compute
The single most important variable is not resolution — it is pixel density on the subject. A 4K camera looking down a 40-metre corridor can put fewer pixels on a person's face than a 1080p camera 12 metres away. Calculate pixels per metre for each critical zone and treat anything under roughly 100 px/m for face-level tasks as a camera upgrade, not a software problem.
Beyond optics, pay attention to wide dynamic range for backlit entrances, IR range for night coverage, and whether the camera's own edge processor can run lightweight detection. Edge-first designs reduce bandwidth and latency, and they keep working when the uplink drops.
Transport, protocols, and bandwidth discipline
RTSP remains the workhorse for pulling streams, ONVIF for device discovery and control, and WebRTC for low-latency viewing in a browser. Codec choice matters: H.265 typically halves bitrate versus H.264 at similar quality, but decode cost is higher, so test it against your GPU budget.
Do the arithmetic early. Forty cameras at 8 Mbps is a continuous 320 Mbps before analytics even start. A common pattern is to analyse a low-resolution substream and only pull the full-quality recording on demand. Keyframe intervals also matter: long GOPs save bandwidth but make seeking and frame-accurate review sluggish.
Metadata, storage, and retention tiers
Video is the expensive part; metadata is the valuable part. Store structured events — timestamp, camera, object class, bounding box, confidence, zone, direction, dwell time — in a relational database such as PostgreSQL (Supabase is a convenient managed option). That gives you fast queries and dashboards. Keep full video in object storage with lifecycle rules: hot for days, warm for weeks, cold archive for the retention period your policy and jurisdiction require.
Thumbnails are a cheap superpower. A 200×200 JPEG per event costs almost nothing and makes review interfaces feel instant.
The model layer
Most production stacks combine several model families rather than one monolith:
- Detection: YOLO-family detectors or RT-DETR for people, vehicles, and object classes.
- Tracking: ByteTrack or BoT-SORT to keep identities stable across frames.
- Classification and embeddings: ResNet or ViT backbones for appearance features used in re-identification.
- Specialists: pose estimation for fall detection, OCR for licence plates, segmentation for precise zone logic.
Version every model, log which version produced which alert, and keep a rollback path. Silent model swaps are how trust in a system dies.
Designing the detection pipeline step by step
Step 1 — Define the events that actually matter
Write the alert list before you write any code. "Detect everything suspicious" is not a specification. Good entries look like: person crosses the perimeter line outward between 22:00 and 06:00; vehicle idles in a fire lane for more than 90 seconds; more than eight people enter Zone C within two minutes.
Step 2 — Choose detection granularity
Decide per event whether you need detection, classification, tracking, or re-identification. Zone intrusion needs detection plus a polygon. Loitering needs tracking plus dwell time. "Find the same person across six cameras" needs re-identification and a much higher bar for legality.
Step 3 — Filter at the edge before the cloud
Send events, not frames. A camera that produces 25 frames per second generates 2.16 million frames a day; the same camera might generate 40 meaningful events. Cascading cheap models first — motion or a small detector — then expensive models only on candidate regions keeps compute sane.
Step 4 — Store structured metadata, not only video
Write every event to the database with enough fields to reconstruct the decision: model version, confidence, bounding box, zone geometry version, and the frame reference. When someone asks "why did this alert fire?", you want an answer, not a shrug.
Step 5 — Close the loop with review and feedback
Alert queues need owners. Mark each alert as true positive, false positive, or undetermined, and feed those labels back into threshold tuning and, eventually, fine-tuning. A system with no feedback loop drifts into noise within months.
Object detection, tracking, and re-identification in practice
Detection is the easy part of the modern stack; identity is where complexity hides. Trackers associate detections across frames using motion prediction and overlap, with appearance embeddings to survive brief occlusions. When two people cross, naive trackers swap identities — and the resulting dwell-time alerts become nonsense.
Practical mitigations:
- Raise the detection confidence threshold for tracking, even if recall drops slightly.
- Use appearance features rather than geometry alone when people are close together.
- Cap track age: after a few seconds without a match, retire the track instead of guessing.
- Test specifically on crowded scenes, not just tidy corridor footage.
Cross-camera re-identification is a different beast. It compares appearance embeddings against a gallery of past observations, which raises two questions teams often skip: how long do you keep the gallery, and who can query it? Gallery hygiene — automatic expiry, deduplication, and strict access control — is as important as the model's accuracy.
A useful rule: only deploy re-identification where the operational value is concrete and documented, such as tracking a suspect vehicle across a campus under an active investigation. Speculative "search everyone, always" designs create legal exposure far beyond their practical benefit.
Anomaly detection and context: beyond simple rules
Rules catch what you already know about. Zones, tripwires, dwell timers, direction filters, and speed thresholds cover the majority of real security use cases and are easy to explain to auditors — which matters more than novelty. Build the rule layer first.
Statistical anomaly detection sits above it. By learning normal activity per camera and per hour-of-week, the system can flag deviations: a loading dock that is normally busy at 09:00 suddenly busy at 02:00, or a footfall pattern that collapses unexpectedly. This is where value shows up in retail shrink, parking enforcement, and industrial safety.
Learned anomaly models — autoencoders, video transformers, and increasingly vision-language models that generate short scene descriptions — add versatility but need careful evaluation. They produce plausible-sounding narratives that can be wrong, so treat their output as a lead, not evidence.
Context fusion is the highest-leverage upgrade. Combine video events with access-control swipes, alarm panel states, shift schedules, and door sensors. A person in a restricted corridor is a curiosity; a person in a restricted corridor at a time when no badge was read is an incident.
Compute, cost, and performance tuning
The economics of analytics are dominated by how many streams you decode and how often you run inference.
Decode offload. Use hardware decoding (NVDEC or equivalent) rather than CPU decoding. It is usually the difference between four and twenty streams per node.
Frame rate reduction. Running inference at 5–8 fps instead of 25 rarely changes event quality for people and vehicles, and it cuts compute proportionally. Fast-moving objects are the exception; handle those with higher-rate analysis only on the relevant cameras.
Quantisation and compilation. INT8 quantisation with TensorRT or OpenVINO typically delivers two to four times the throughput of unoptimised FP32 inference with modest accuracy loss — verify on your own validation set rather than trusting benchmarks.
Cascades and batching. Run a small model everywhere and a large model only on candidate regions. Batch inference requests on a GPU; unbatched single-frame calls waste most of the hardware.
Autoscaling. If you run in the cloud, scale inference on queue depth rather than a fixed instance count. Preemptible capacity is acceptable for retrospective indexing, never for live alerting.
Track four metrics weekly: false alerts per camera per day, missed-event rate on a labelled test set, end-to-end alert latency, and GPU utilisation. If false alerts per camera per day rises, users stop reading the queue — and the entire investment silently evaporates.
Privacy, compliance, and responsible deployment
Analytics that identify people sit squarely in regulated territory. Treat compliance as a design input, not a launch checklist.
- Purpose limitation. Collect and analyse only what a documented purpose requires. Blanket recording of audio is restricted or banned in many jurisdictions.
- Notice and transparency. Signage, internal policies, and, for employees, consultation where required.
- Retention limits. Automate deletion. Metadata that lives forever is a liability, not an asset.
- Access control. Role-based permissions, immutable audit logs of every query and export.
- Minimisation by design. Privacy masking for neighbouring properties, blurring for public areas, and disabling face or plate recognition unless there is a specific legal basis.
- Biometrics. Facial recognition and similar techniques usually attract a higher legal bar than general motion detection. Many organisations can meet their objectives with non-biometric approaches.
- Location of processing. Some sectors require on-premises inference. Confirm this before choosing a cloud-first architecture.
Document a data protection impact assessment early. It forces the conversations — retention, access, legal basis — that otherwise surface after an incident.
Common mistakes and how to avoid them
- Tuning in the lab only. Daylight test clips hide the night, rain, and headlight problems that define real performance.
- Ignoring camera hygiene. Dirty domes, spider webs, and drifting PTZ presets generate more false alerts than model error ever will.
- Alert flooding. If an alert fires 200 times a day, it is not an alert, it is noise. Use thresholds, cooldowns, and zone exclusions aggressively.
- No single owner for the queue. Assign accountability, or the queue becomes an archaeological site.
- Unsynced clocks. Without reliable NTP, correlating video with access logs becomes guesswork.
- Storing everything forever. Cost grows linearly; value does not.
- One model to rule them all. Different scenes need different thresholds. Per-camera configuration is normal, not a failure.
- Skipping the feedback loop. Systems improve because humans label mistakes, not because vendors ship new weights.
A 30-day pilot plan
Week 1 — Baseline. Choose 8–12 cameras covering your highest-value zones. Fix camera hygiene, verify NTP, and record a week of normal activity without analytics to establish a baseline.
Week 2 — Deploy the rule layer. Zones, tripwires, dwell timers, and direction filters. Keep the model simple and the configuration explainable.
Week 3 — Tune with humans. Have operators label every alert. Adjust thresholds per camera, add exclusions, and measure false alerts per camera per day.
Week 4 — Add context and decide. Fuse access-control or alarm data, then compare measured detection performance against the baseline. Decide to expand, hold, or stop — with numbers rather than enthusiasm.
Frequently asked questions
How accurate is AI video analytics in practice? For well-framed people and vehicle detection in daylight, modern detectors exceed 90% recall on typical scenes. Night, rain, and crowds degrade this substantially. Always evaluate on footage from your own cameras.
Do I need to replace my existing CCTV? Usually not. If cameras meet pixel-density targets and stream via RTSP or ONVIF, they can feed an analytics platform. Upgrade only the cameras covering critical zones.
Should inference run on the camera, on-premises, or in the cloud? Edge for latency and bandwidth, on-premises for privacy-sensitive sites, cloud for retrospective indexing and elasticity. Many deployments use all three.
How much compute does one stream need? Roughly one mid-range GPU can handle 8–20 streams at reduced frame rate with an optimised small model. Specialist models or high frame rates reduce that considerably.
Can analytics replace security staff? No. It replaces staring at monitors. Humans still decide, respond, and take responsibility for outcomes.
What is the fastest way to reduce false alerts? Tighten zones, add cooldowns, raise confidence thresholds, and fix camera framing. Model upgrades come last, not first.
Is facial recognition required for access control? Almost never. Badge, PIN, or non-biometric tailgating detection usually achieves the same operational goal with far less regulatory friction.
How do I know the system is working a year later? Track false alerts per camera per day, alert acknowledgement rates, and missed events on a retained labelled test set. Review quarterly, not annually.



