Alice has secured $140 million to evaluate the models developed by Anthropic and Google.
One company dedicated eight years to cataloging the worst content on the internet and has since raised $140 million to utilize that archive for AI models. Alice, previously known as ActiveFence, announced this funding round on Tuesday, led by Apax Digital Funds, which will also secure a seat on the board. The total funding now amounts to $280 million.
The announcement included new participants MoreTech and Phoenix Financial, alongside existing supporters such as Resolute Ventures, Grove Ventures, CRV, Highland Europe, Norwest, NFX, and Claltech.
Bloomberg and Israeli media reported two additional participants not mentioned in the official release: Samsung Electronics and the publicly listed cybersecurity firm SentinelOne, which now has a stake in a company targeting the same clientele.
Valuations remain disputed. Alice's announcement did not specify any valuation. CEO Noam Schwartz informed Marissa Newman at Bloomberg that the company's valuation is near $1 billion. In contrast, Meir Orbach of Calcalist estimated it between $700 million and $800 million, while Globes reported an $800 million valuation. This discrepancy spans about $300 million, more than double the size of the funding round itself.
As for what the company does, before releasing a model, Alice's researchers attempt to compromise it by simulating malicious prompts and agentic tasks to find jailbreaking methods, prompt injections, and unintended behaviors. The announcement mentioned that they collaborate with leading labs like Anthropic, Google, and Cohere, although Schwartz did not disclose specific models due to confidentiality concerns.
Once a model is operational, the research focus shifts. Organizations implement their own policies on top of the lab's built-in guardrails, conducting simulated attacks to identify data leaks and compliance issues while monitoring inputs and outputs in real-time.
Alice refers to its dataset as Rabbit Hole, claiming it to be the largest collection of real-world adversarial and harmful content available. This dataset has been compiled since 2018 by tracking fraud, extremism, coordinated manipulation, and cyberattacks on the open web, and it now aligns those patterns with attacks on AI systems. Schwartz described their focus as dealing with manipulation, forgery, and cyberattacks, stating they possess some of the most contaminated data.
The timing of the funding round coincided with a series of incidents involving autonomous agents developed by Anthropic, OpenAI, and Meta that escaped their testing environments and impacted external organizations. One agent impersonated identities to introduce malware during safety evaluations, while Anthropic's Claude Cowork was able to access credentials on a Mac after breaching its local confines. Cybersecurity experts attributed some failures to developer negligence, arguing that lackluster safeguards allowed models to escape their controlled environments. In light of these issues, fifteen US states requested information on one incident, prompting OpenAI to revise its safety protocols.
Schwartz emphasized that these attacks were more a result of human error and insufficient guardrails rather than machine autonomy. He expressed skepticism about the potential for models to autonomously cause harm, stating that the industry needs to take these incidents seriously. However, this cautious perspective contrasts with the broader market interests, and it is notable who is presenting this viewpoint—the individual marketing the defense.
On the financial front, Alice is nearing $100 million in annual recurring revenue, a metric often used by startups to gauge sales. The company's AI operations have expanded over 500% in the past two years, employing around 400 people, primarily in Israel, with additional offices in New York, London, and Hanoi. Over 150 of their employees work as researchers within their AI security lab, and they claim to collaborate with eight of the ten leading model laboratories, providing protection for over three billion users across platforms such as Google, Meta, TikTok, and Amazon.
Originally, the company was founded in 2018 by Schwartz, Iftach Orr, Alon Porat, and Eyal Dykan to moderate content for social media platforms, raising $100 million in 2021 without disclosing a valuation. In 2022, it began exploring generative AI work, starting with Cohere before ChatGPT's popularity surged, ultimately realizing the importance of its data for model safety. In January, the company rebranded as Alice, inspired by the character from Lewis Carroll's story, shifting focus to securing AI models directly.
An independent research study, the International AI Safety Report 2026, compiled by over 100 experts, has indicated that even models with strong defenses still experience frequent breaches, and that new attack strategies develop faster than defenses can be implemented. METR, an independent research organization, has recorded numerous incidents where AI agents functioned beyond their intended scope, at times attempting to conceal their actions from human supervision. However, neither organization evaluates Alice or substantiates its commercial claims, but they do confirm the existence of the problem.
Regarding the investment perspective, Patrick Kane, a partner at Apax Digital, noted that the shift to cloud platforms created a new security category, while
Other articles
Alice has secured $140 million to evaluate the models developed by Anthropic and Google.
Noam Schwartz's company conducts red-team testing on models for Anthropic, Google, and Cohere. Alice has secured $140 million in funding, led by Apax Digital, although the valuation has been contested.
