Video Analysis with Local AI
What This Article Covers
- How to prepare videos for multimodal AI models.
- The steps required from file to description.
- Which hardware and software suit local video analysis.
- How to extract frames, create chunks, and feed them to models.
- Data privacy, performance, and common pitfalls.
Introduction: Video Analysis with Local AI
Videos contain more information than individual images. Motion, audio, temporal flow, and context together form a complex data source. Multimodal AI models can describe individual frames, recognize objects, and answer questions about scenes. Local video analysis makes this possible without sending sensitive recordings to external providers.
The challenge lies in processing. A single video can contain thousands of frames. Large language models cannot see all frames at once. Anyone wanting to analyze videos locally needs a strategy to select the right frames, compress them, and feed them meaningfully to models.
Why Do You Need Local Video Analysis?
Surveillance cameras, training videos, documentation, and content production generate large volumes of video. Local analysis protects privacy and avoids expensive cloud fees. Video analysis can be used to:
- Summarize scenes,
- Count objects or people,
- Detect anomalies,
- Describe content for accessibility,
- Catalog videos for archives.
Control over the data remains within your own network. This is especially relevant for companies, schools, and government agencies.
Video Analysis Explained
The typical workflow has several steps:
- Load video: The video exists as a file or stream.
- Extract frames: Individual images are read from the video at regular intervals.
- Select frames: Not every image is relevant. Keyframes, scene changes, or motion areas reduce the data volume.
- Prepare images: Adjust size, format, and quality.
- Run model: A multimodal model describes frames or answers questions.
- Aggregate results: Combine multiple results into an overall picture.
Key concepts:
- Frames per second (fps): Number of images per second.
- Keyframe: An individual frame that is not calculated only from differences.
- Scene detection: Automatic recognition of cuts or scene changes.
- Temporal context: Time-based relationships between frames.
- Embedding: Numerical representation of image content for search.
Who Is Video Analysis For?
- Security officers who need to review camera feeds.
- Teachers and trainers who want to structure video content.
- Content teams who catalog material.
- Developers who integrate multimodal models into applications.
- Privacy-conscious users who don’t want to upload data.
Key Terms in Video Analysis
- Multimodal model: AI that processes text and images together.
- Frame extraction: Reading individual images from a video.
- Optical flow: Analysis of motion between frames.
- Video chunking: Dividing into smaller time segments.
- Scene graph: Structure describing objects and relationships in scenes.
- Transcript: Text representation of spoken content, often from audio.
Technical Requirements
Hardware
Video analysis is computationally intensive. At minimum:
- A GPU with 12 GB VRAM for smaller multimodal models.
- 16 GB RAM, preferably 32 GB.
- Fast NVMe storage for large video files.
- Multi-core CPU for preprocessing and frame extraction.
For larger models like Qwen2-VL, InternVL, or LLaVA-1.6, 24 GB VRAM or more is recommended. If you only analyze occasionally, you can work with an 8 GB model and fall back to CPU processing.
Software
- OpenCV: Frame extraction and preparation.
- FFmpeg: Conversion, resizing, and scene detection.
- Python: Pipeline scripting.
- Ollama or vLLM: Running multimodal models.
- LLaVA, BakLLaVA, Qwen2-VL: Multimodal models.
Practical Example: Extract and Analyze Frames
1. Extract frames with FFmpeg
ffmpeg -i video.mp4 -vf "fps=1/5,scale=512:-1" -q:v 2 frames/%04d.jpg
This command generates one image every five seconds at 512 pixels wide.
2. Detect scene changes
FFmpeg also supports scene detection:
ffmpeg -i video.mp4 -filter_complex "select=gt(scene\,0.3),scale=512:-1" -vsync vfr frames/scene_%03d.jpg
The threshold 0.3 determines how much a frame must change to be considered a new scene.
3. Python pipeline for batch processing
import ollama
import os
from glob import glob
frames = sorted(glob('frames/*.jpg'))
descriptions = []
for image in frames:
response = ollama.chat(
model='llava',
messages=[{
'role': 'user',
'content': 'Describe briefly what you see in this image.',
'images': [image]
}]
)
descriptions.append(response['message']['content'])
print('\n'.join(descriptions))
4. Summarize across all frames
summary = ollama.chat(
model='llama3.1',
messages=[{
'role': 'user',
'content': 'Summarize the following scene descriptions into a brief video overview: ' + ' | '.join(descriptions)
}]
)
print(summary['message']['content'])
Strategies for Reducing Data Volume
Videos often contain many redundant frames. Without reduction, the model becomes overwhelmed. Common approaches:
- Temporal downsampling: Use only every nth frame.
- Scene detection: Only frames at scene changes.
- Motion detection: Only consider moving areas.
- Clustering: Group similar frames together.
- Use keyframes: Leverage frames already marked by the codec.
A combination of scene detection and temporal reduction typically offers the best balance between information content and computational cost.
Including Audio
Many videos contain important information in the audio track. Speech, sound effects, and music contribute to interpretation. Combining video analysis with speech-to-text enables:
- Transcription of spoken content,
- Synchronization of image and speech,
- Searching for specific terms in the video,
- Summarization combining both image and audio.
For speech recognition alone, whisper.cpp or faster-whisper can run locally. Results can be merged with frame descriptions.
Common Pitfalls in Video Analysis
- Too many frames: Analyzing every frame is inefficient and expensive.
- Resolution too high: Large images consume a lot of VRAM. Scale to reasonable sizes.
- Model too small: Smaller models miss details or contextual relationships.
- No image captions: Answers without context are often inaccurate.
- Audio overlooked: Many videos only make sense with audio.
- Personal data: Respect data protection and retention requirements.
Further Links and Resources
FAQ: Video Analysis with Local AI
How many frames should I analyze? It depends on the content. For summaries, one or two frames per second are often sufficient. For detailed analysis, use scene changes.
Do I need a GPU? Yes, if processing needs to complete in reasonable time. CPU modes are slow for occasional testing.
Which models are suitable? LLaVA, BakLLaVA, Qwen2-VL, and InternVL are good open-source options.
Can I analyze live streams? Yes, with proper buffering and reduction. Latency increases with model size.
What about privacy-sensitive videos? Local processing is required. Avoid uploads and enforce strict storage and deletion policies.
Sources and Further Reading
- LLaVA: https://llava-vl.github.io/
- OpenCV: https://opencv.org/
- FFmpeg: https://ffmpeg.org/
- Ollama: https://ollama.com/
Summary: Video Analysis with Local AI
Local video analysis requires solid preparation. Extracting frames, filtering scenes, scaling images, and selecting an appropriate multimodal model are the key steps. When audio and video are combined, more complete results emerge. Hardware choices, data privacy, and data reduction determine whether your project runs efficiently and remains legally compliant.


