Why video serves as the initial frontier as AI begins to interpret the physical world.
TL;DR: There are already hundreds of millions of cameras installed, but most footage requires human intervention to determine what is important. Lumana, founded by former Intel computer vision experts, processes over a billion images daily from more than 50,000 cameras using its VIA-1 model, which learns the usual activity of each camera and flags any anomalies. The company filters data locally before sending anything to the cloud, adhering to the "filter before you spend" principle. Video may be an optimal entry point for physical AI due to existing infrastructure.
For instance, a camera monitoring a loading dock could capture twelve hours of footage of trucks, workers, and packages, which is largely unviewed until something goes wrong, such as a package disappearing or an accident. This has been a significant limitation of video surveillance, as while cameras have transitioned to digital, the footage still heavily relies on human insight to determine what to search for. AI is beginning to enhance video footage, making it more understandable and searchable during real-time events.
Axis Communications forecasts that by the end of 2025, 562 million surveillance cameras will be in use globally, excluding China. An increasing number of these cameras will feature built-in intelligence, as approximately two-thirds of those shipped in 2024 are expected to include deep-learning analytics, according to the company's research.
Companies developing what is termed physical AI find that much of the required infrastructure is readily available. The key opportunity lies in enhancing the utility of existing cameras by teaching software to interpret the actions occurring in front of them. Lumana is among these companies, with its founders having substantial experience in computer vision at Intel. CEO Sagi Ben Moshe previously led Intel's RealSense division, and CTO Ofir Mulla worked on the design for its 3D and LiDAR cameras.
The California-based startup secured $40 million in July 2025 in a Series A funding round led by Wing Venture Capital, with contributions from Norwest Venture Partners and S Capital, bringing its total funding to $64 million. By December, the company reported over 50,000 cameras connected to its platform, serving Fortune 500 clients across the United States.
Lumana states that its AI video surveillance systems now process over a billion images daily across more than 50,000 cameras. However, for customers, the initial changes are often more mundane than the staggering figure implies.
Initially, operators are encouraged to stop continuously monitoring multiple screens. For years, security control rooms have typically featured walls of video feeds, with personnel expected to notice anomalies. According to Ofir Mulla, Lumana's CTO, teams spend the first month after implementing the technology addressing common and routine issues, such as offline cameras, inadequate coverage, and determining who should receive alerts.
As these tasks are resolved, operators can redirect their focus from watching all the feeds to examining the incidents flagged for further investigation. "Instead of dedicating time to observing passive video walls and locating unusual occurrences, operators can concentrate on events that require scrutiny,” Mulla explained in an interview.
The effectiveness of this approach is heavily context-dependent, something traditional motion detection systems have struggled to manage. For example, an individual beside a warehouse door may seem normal at 2 PM but might warrant investigation at 2 AM. A delivery truck parked at a loading dock for 20 minutes could be expected, while the same truck staying for three hours might not be.
Lumana's VIA-1 model is designed to understand the specific environment of each camera rather than applying a uniform definition of normalcy. The company asserts this can decrease false alerts by up to 90% in comparison to traditional motion detection and rule-based systems, although Mulla specifies that this statistic reflects the higher end of Lumana’s findings rather than a guaranteed outcome for all customers.
The challenge becomes more complex when the environment changes. What a camera perceives as normal can shift rapidly. A warehouse may rearrange its layout over a weekend, or during the holiday season, retailers may face increased customer traffic, or a factory may introduce a night shift with numerous employees in previously quiet areas. Activities that might have triggered alerts yesterday might be entirely routine today.
Mulla indicated that VIA-1 can adapt its understanding as environments evolve, supported by operator feedback when significant changes occur. He refrained from providing a fixed timeframe for these adjustments as it varies based on the nature and scale of the change.
An effective AI video system must recognize when new patterns become routine while also detecting activities that require further investigation. Since the duration for adaptation hinges on the changes made, operator input is essential, particularly when substantial modifications occur at a location. In reality, the system evolves concurrently with a dynamic environment.
Additionally, improving the accessibility and comprehensibility of this footage brings up privacy concerns, especially in workplaces where individuals are frequently recorded. Lumana allows clients to configure retention and access protocols and disable features like facial and gender recognition based on local regulations. However, with enhanced camera capabilities, companies must carefully consider
Other articles
Why video serves as the initial frontier as AI begins to interpret the physical world.
With 562 million surveillance cameras already in place globally and two-thirds now equipped with deep-learning analytics, Lumana believes that enhancing existing video feeds to be searchable and intelligent is the quickest way for physical AI to achieve scale.
