The video-review problem
Your cameras record everything — and understand nothing.
Traditional video systems are digital recorders that alert on rigid motion rules. When something actually happens, someone scrubs through hours of footage by hand. VizAIo understands what happened — and lets you just ask.
Rule-based VMS & manual review
- ✕Hours of manual scrubbing. Finding one incident means an operator watching footage frame by frame across dozens of cameras.
- ✕Motion, not meaning. Pixel and motion rules fire on anything that moves — they can't tell a delivery from an intrusion.
- ✕Rigid and single-purpose. Fixed rules cover one security scenario at a time and miss anything they weren't configured for.
- ✕Expensive to scale. Per-camera licences, dedicated hardware and per-stream pricing make broad coverage cost-prohibitive.
Ask, and get frame-referenced answers
- ✓Seconds, not hours. A plain-English query jumps straight to the exact moment — a 2-hour review becomes a 5-second search.
- ✓Understands the scene. Gemini Vision profiling plus semantic search knows people, objects, actions and anomalies — not just movement.
- ✓Multi-domain by design. One platform covers retail, EHS, healthcare and security — no per-scenario rebuild.
- ✓Runs on what you have. Works on existing IP cameras at $0.00045 per operation — no hardware swap.
Enterprise-grade unit economics
From hours of footage to
a five-second answer
%
Time saved — a 2-hour manual review becomes a 5-second natural-language query
$0.00045
Per AI operation — optimized frame sampling makes video intelligence fractions of a penny
s
Intelligence loop — live frames analysed every 3 seconds for near real-time awareness
%
Gross operating margin — dynamic model routing keeps infrastructure cost low at scale
How it works
One streaming pipeline —
from raw frames to a forensic
answer
VizAIo ingests live and archived video, profiles every scene with vision AI,
indexes it for semantic search, then reasons over your query and streams the
answer back — near real time, on the cameras you already run. Tap any stage.
Live feeds and archives, one pipeline
IngestAgent pulls live CCTV over RTSP (TCP :8554) through an EC2 MediaMTX relay, and archived MP4s (up to 2GB) via multipart uploads to S3 — clipping one frame every 3 seconds with FFmpeg.
AI understands each scene
ProfilerAgent routes frame buffers through Gemini Vision (describe_segment) with a strict caption prompt to classify people, objects, actions, layouts and exceptions — writing time-synchronised JSON metadata and frames to an S3 data lake.
Every event becomes searchable
LinkAgent embeds content into 3072-dimensional vectors (gemini-embedding-001) and indexes them in ChromaDB with HNSW cosine similarity and a greedy time-diversity filter — no manual tagging.
Plain-English forensic reasoning
A natural-language query runs an HNSW similarity search to retrieve the right frames, then ReasonAgent interprets the spatial-temporal context with forensic system instructions to explain what actually happened.
Answers arrive as they form
SynthAgent serialises the reasoning into live Server-Sent Events over FastAPI, while OrchestratorAgent manages async pipeline state in PostgreSQL — no third-party storage layer required.
Alerts, answers and evidence
Out comes a streamed answer with frame timestamps, 7-day pre-signed S3 asset references and confidence metrics — plus instant alerts and court-ready incident timelines.
See it in action
Stop watching. Start
knowing — live.
Watch a live feed become searchable intelligence: frame ingestion, Gemini
Vision scene profiling, semantic retrieval, and a streamed forensic answer with
the exact frames and timestamps behind it.
Why VizAIo
"Isn't this just video
analytics?" No — and here's
exactly why.
Watch a live feed become searchable intelligence: frame ingestion, Gemini
Vision scene profiling, semantic retrieval, and a streamed forensic answer with
the exact frames and timestamps behind it.
Understanding, not just motion
Natural-language search, not tag-and-scrub
Live + archived, not one or the other
Your existing cameras, not new hardware
Petabyte-scale economics, not per-camera pricing
Understanding, not just motion
Rule-based systems fire on pixel and motion changes. They flag that something moved, never what it was — flooding operators with false alerts and missing the events that matter.
ProfilerAgent routes frames through Gemini Vision to classify people, objects, actions and layouts, then reasons over spatial-temporal context — so VizAIo knows a delivery from an intrusion.
Natural-language forensic search, not tag-and-scrub
Finding an event means either pre-tagging everything up front or scrubbing hours of timeline by hand — slow, manual, and only as good as the tags someone remembered to add.
Describe what you're looking for — "person in a red jacket near the loading dock after 11pm" — and HNSW vector search returns the exact frames. No tagging, no indexing, no scrubbing.
Live and archived, in one pipeline
Most tools do one or the other — live monitoring or uploaded-clip analysis — forcing separate systems and fragmented workflows.
One pipeline ingests live RTSP feeds via an EC2 MediaMTX relay and archived MP4 uploads to S3 — real-time awareness and deep forensic history in the same searchable index.
Works on the cameras you already have
Generic CV platforms often need specific hardware, edge devices or camera replacements — a costly, disruptive rollout before you see any value.
VizAIo runs on your existing IP and RTSP cameras with no hardware swap — plug-and-play ingestion means you can be searching real footage in a single working session.
Petabyte-scale economics, not per-camera pricing
Fixed licences plus hardware, or per-camera and per-stream pricing, make broad, always-on AI coverage prohibitively expensive.
Optimized frame sampling and dynamic model routing (Gemini Flash ↔ Pro) drive the core AI cost to $0.00045 per operation — enterprise video intelligence with a 95% gross margin built in.
What it does
Three capabilities, one
private pipeline
Intelligent Video Analysis & Extraction
Turns raw footage into structured visual intelligence.
- Securely ingests live RTSP feeds, CCTV streams and archived uploads across the enterprise.
- Samples optimized frames every 3 seconds and identifies people, objects, actions and anomalies.
Conversational Video Intelligence
Interrogate hours of footage in plain English.
- Ask natural-language questions like "show all unauthorized access after 9 PM."
- Semantic forensic search locates key moments in seconds and summarizes incidents.
Automated Detection & Alerting
Catch risks before they escalate into incidents.
- Behavioral analysis flags suspicious movement, unusual dwell time and abnormal activity.
- Detects intrusion, loitering, falls and unauthorized access, with instant SSE alerts.
Built for enterprise trust
Cost-efficient AI under the
hood. Plugs into everything
you run.
Enterprise video intelligence at a fraction of a penny
VizAIo pairs vision models with a semantic index and a streaming backend — with dynamic routing that keeps the cost per operation tiny at scale.
Built to connect to everything you run
VizAIo doesn't operate in isolation — it feeds into your systems and Prajna AI's ecosystem, and adapts as your environment changes.
Feeds video intelligence into PrajnaAI's Data Fabric for enterprise-wide decision intelligence.
Connects with CCTV platforms, monitoring dashboards, security systems and third-party APIs.
Continuously adapts to new environments, workflows and evolving operational scenarios.
Stop watching. Start knowing.
Your team is spending thousands of hours reviewing video that VizAIo can analyse in seconds — at $0.00045 per operation, with 95% gross margin built into the architecture.