[2608.08075] Search over the Visual World: Persistent Visual Memory, Layered Indexes, and Source-Grounded Evidence
Skip to main content
Search arXiv
Press Enter to search · Advanced search
-->
Computer Science > Information Retrieval
arXiv:2608.08075 (cs)
[Submitted on 8 Aug 2026]
Title:Search over the Visual World: Persistent Visual Memory, Layered Indexes, and Source-Grounded Evidence
Authors:Sankalp Nagaonkar, Rohit Garg, Ankit Raj, Ashish Choithani, Ashutosh Trivedi<br>View a PDF of the paper titled Search over the Visual World: Persistent Visual Memory, Layered Indexes, and Source-Grounded Evidence, by Sankalp Nagaonkar and 4 other authors
View PDF<br>HTML (experimental)
Abstract:Most video-retrieval systems assume a bounded corpus and return ranked files or timestamps. Agents operating over cameras, screens, streams, and archives face a different systems problem: observations arrive continuously; models interpret them at different temporal granularities; context must be selected without replaying the complete visual record; and results must stay connected to inspectable source evidence. We argue that search over such a corpus is an infrastructure problem that cannot be reduced to ranking video files. We develop a conceptual and formal model of search over the visual world built on analyzer-defined scenes, persistent understanding artifacts, visual memory as coexisting scene spaces over shared source time, and capability-declared indexes, distinguishing memory (everything retained), context (what is selected for a task), and evidence (the source intervals that ground it). The VideoDB data format (VDB) realizes this model in production, exposed through a typed search surface spanning planned retrieval, stateful investigation, direct access, and grounded synthesis. We contrast this model-agnostic infrastructure, where segmentation, sampling, model choice, embeddings, and ranking are system decisions and live streams are first-class sources, with video-native foundation models offered as fixed APIs. In a semantic-retrieval comparison against a commercial video-native engine spanning 9,800+ queries over four public datasets, a pipeline of general-purpose components achieves higher macro-averaged Recall@1/@3/@10 (73.09/83.39/91.20 versus 65.75/77.13/89.10), while the baseline is higher at Recall@50 (96.42 versus 96.07). Retrieval quality over the visual world is today governed more by system design than by video-specific pretraining, and visual-memory infrastructure can deliver it while keeping playable, source-grounded evidence first-class.
Comments:<br>33 pages, 5 figures, 17 tables. Technical report. Benchmark configurations and reproduction instructions: this https URL
Subjects:
Information Retrieval (cs.IR); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
Cite as:<br>arXiv:2608.08075 [cs.IR]
(or<br>arXiv:2608.08075v1 [cs.IR] for this version)
https://doi.org/10.48550/arXiv.2608.08075
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)
Submission history<br>From: Ashish Choithani [view email]<br>[v1]<br>Sat, 8 Aug 2026 11:46:54 UTC (53 KB)
Full-text links:<br>Access Paper:
View a PDF of the paper titled Search over the Visual World: Persistent Visual Memory, Layered Indexes, and Source-Grounded Evidence, by Sankalp Nagaonkar and 4 other authors<br>View PDF<br>HTML (experimental)<br>TeX Source
view license
Additional Features
Audio Summary
Current browse context:
cs.IR
next >
new<br>recent<br>| 2026-08
Change to browse by:
cs<br>cs.CV<br>cs.MM
References & Citations
NASA ADS<br>Google Scholar
Semantic Scholar
export BibTeX citation<br>Loading...
BibTeX formatted citation
×
loading...
Data provided by:
Bookmark
Bibliographic Tools
Bibliographic and Citation Tools
Bibliographic Explorer Toggle
Bibliographic Explorer (What is the Explorer?)
Connected Papers Toggle
Connected Papers (What is Connected Papers?)
Litmaps Toggle
Litmaps (What is Litmaps?)
scite.ai Toggle
scite Smart Citations (What are Smart Citations?)
Code, Data, Media
Code, Data and Media Associated with this Article
alphaXiv Toggle
alphaXiv (What is alphaXiv?)
Links to Code Toggle
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub Toggle
DagsHub (What is DagsHub?)
GotitPub Toggle
Gotit.pub (What is GotitPub?)
Huggingface Toggle
Hugging Face (What is Huggingface?)
ScienceCast Toggle
ScienceCast (What is ScienceCast?)
Demos
Demos
Replicate Toggle
Replicate (What is Replicate?)
Spaces Toggle
Hugging Face Spaces (What is Spaces?)
Spaces Toggle
TXYZ.AI (What is TXYZ.AI?)
Related Papers
Recommenders and Search Tools
Link to Influence Flower
Influence Flower (What are Influence Flowers?)
Core recommender toggle
CORE Recommender (What is CORE?)
Author
Venue
Institution
Topic
About arXivLabs
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on...