Why video serves as the initial frontier as AI begins to comprehend the physical world

Why video serves as the initial frontier as AI begins to comprehend the physical world

      TL;DR: Hundreds of millions of cameras are already in place, yet most footage still needs human oversight to identify key events. Lumana, created by former Intel computer vision experts, processes over a billion images daily from more than 50,000 cameras using its VIA-1 model, which learns the typical behavior of each camera and identifies deviations. The company filters footage locally before uploading to the cloud, adhering to the principle of "filter before you spend." Video may be an ideal entry point for physical AI due to the existing infrastructure.

      A camera monitoring a loading dock may capture twelve hours of footage showing trucks arriving, workers moving about, and boxes departing. On most days, there’s little reason to review this footage, which only becomes valuable when incidents like a missing package or an accident occur, or when someone needs to trace back events.

      This limitation has plagued video surveillance for years; while cameras have been digital for some time, the footage often requires a person to know what to search for. AI is now making such footage more comprehensible and searchable in real time.

      Axis Communications predicts that by the end of 2025, 562 million surveillance cameras will be operational worldwide (excluding China). An increasing number of these cameras will come with built-in intelligence, with around two-thirds being equipped with deep-learning analytics in 2024.

      For companies focused on developing what's known as physical AI, much of the necessary infrastructure is already in place. The goal is to enhance the utility of existing cameras by training software to understand the activities occurring in front of them.

      Lumana is a company investing in this evolution, founded by individuals with extensive experience in computer vision from Intel. The CEO, Sagi Ben Moshe, previously led Intel's RealSense division, while CTO Ofir Mulla contributed to the design of its 3D and LiDAR cameras.

      In July 2025, the California-based startup secured $40 million in a Series A funding round led by Wing Venture Capital, with additional support from Norwest Venture Partners and S Capital, bringing its total funding to $64 million. By December, it had over 50,000 cameras linked to its platform, serving Fortune 500 clients across the U.S.

      The company claims its AI video surveillance systems process over a billion images daily from more than 50,000 cameras. However, the initial changes for clients are often more mundane than the numbers imply.

      First, stop watching the screens.

      For years, security control rooms have featured walls of video feeds requiring personnel to monitor for anomalies actively.

      According to Lumana CTO Ofir Mulla, teams spend the initial month post-deployment addressing typical problems, identifying offline cameras, inconsistent coverage, and alert assignment.

      Once those issues are resolved, operators can shift their focus from passive monitoring of every feed to investigating flagged events. “Instead of watching passive video walls and trying to spot something unusual, operators can concentrate on events that warrant further examination,” Mulla stated in an interview.

      Effective operation heavily relies on context, an area where traditional motion detection and fixed rules often fall short. For example, a person standing next to a warehouse door may be ordinary at 2 PM but suspicious at 2 AM. Similarly, a delivery truck parked at a loading dock for 20 minutes might be normal, but if it remains for three hours, it could raise alarms.

      Lumana’s VIA-1 model adapts its understanding based on each camera's environment rather than applying a uniform standard across all. The company claims this can reduce false alerts by up to 90% compared to older motion detection systems, though Mulla underscores that this figure represents the upper limit observed rather than a guarantee for every customer.

      The challenge intensifies when the environment itself shifts.

      What occurs when normal changes?

      A camera's definition of "normal" can quickly evolve. A warehouse might reorganize its layout over a weekend, retailers may face a Christmas rush, or factories may introduce night shifts with many employees arriving in previously empty areas. Activities that triggered alerts in the past could become routine.

      Mulla notes that VIA-1 can adjust its understanding of a camera as circumstances change, utilizing operator feedback when significant changes occur. He refrained from providing a specific timeframe for these adjustments, as it varies based on the nature and extent of the changes.

      An effective AI video system must recognize a new shift pattern as standard while still identifying unusual activities that may require attention. Given that adjustment time hinges on what has changed, operator input remains crucial, particularly after considerable changes at a site. Essentially, the system continually learns in a dynamic environment.

      Making footage more searchable and understandable also raises privacy concerns, especially in workplaces where individuals are frequently recorded. Lumana enables customers to establish retention and access policies, and to deactivate features like facial and gender recognition depending on local regulations. As existing cameras grow more capable, companies must also ponder what data to collect, who can search it, and the duration for which this information should be retained.

      Currently, much

Why video serves as the initial frontier as AI begins to comprehend the physical world

Other articles

Nvidia has announced its acquisition of Hugging Face for $12.93 billion and assures that it will maintain its openness. Nvidia has announced its acquisition of Hugging Face for $12.93 billion and assures that it will maintain its openness. Nvidia has announced a $12.93 billion purchase of Hugging Face, giving the platform a valuation of approximately 86 times its revenue. Jensen Huang stated that it will continue to be open across various models, frameworks, cloud services, and computing platforms. New Jersey has recently authorized affordable plug-in solar panels for balconies. New Jersey has recently authorized affordable plug-in solar panels for balconies. New Jersey has prohibited landlords and municipalities from preventing plug-in solar installations. In 2024, Germany granted this right to tenants, and just last week, Britain legalized such systems. Best alternatives to Rippling for mid-market businesses. Best alternatives to Rippling for mid-market businesses. A ranking of 12 HR platforms for mid-market purchasers seeking alternatives to Rippling's IT-focused package. This review includes HiBob, Paylocity, Gusto, BambooHR, Deel, Workday, ADP, Paychex, Personio, Justworks, Paycom, and Remote, detailing pricing, advantages, and issues noted by reviewers. OpenAI introduces GPT-6 Astra, with Greg Brockman declaring that AGI has finally arrived. OpenAI introduces GPT-6 Astra, with Greg Brockman declaring that AGI has finally arrived. The president of OpenAI states that Astra meets the criteria for AGI. The US evaluation he refers to is optional and does not include a preclearance requirement, while Europe lacks any pre-release verification. A Tesla-like display can update an older Silverado, but the challenging aspect is ensuring all other components function properly. A Tesla-like display can update an older Silverado, but the challenging aspect is ensuring all other components function properly. Pickup trucks have the potential to last for many years, but their infotainment systems become outdated quickly. Merge Screens provides Tesla-like Android displays for Silverados from 2007 to 2026, but the significant challenge lies in integrating them with steering-wheel controls, cameras, climate systems, and original audio systems. The top 10 autonomous penetration testing tools, ranked based on exploit validation (2026). The top 10 autonomous penetration testing tools, ranked based on exploit validation (2026). Ten autonomous pentesting platforms were evaluated based on proof-of-exploit, attack-chain depth, and coverage. Astra Security comes out on top for web and API exploitation, featuring a protected validator that re-exploits each finding before it enters your queue.

Why video serves as the initial frontier as AI begins to comprehend the physical world

With 562 million surveillance cameras currently in place globally and two-thirds of them now equipped with deep-learning analytics, Lumana is wagering that transforming existing video feeds into searchable and intelligent sources is the quickest route for physical AI to achieve scalability.