Neel Somani on Why AI Labs Are Looking Beyond the Web for Future Training Data

Facebook
X
WhatsApp
Table of Contents
Neel Somani on Why AI Labs Are Looking Beyond the Web for Future Training Data

For years, building more capable AI models followed a simple formula: gather more data.

That approach is becoming harder and harder to sustain.

The supply of high-quality public text is finite, copyright disputes have made large-scale web scraping more legally complicated, and the internet itself is increasingly filled with AI-generated content that can reduce the quality of future training data.

As AI labs look beyond traditional web data, the focus is shifting to a new question: where will the next generation of training and evaluation data come from?

AutomataBench, a benchmark created by Neel Somani, a former quantitative researcher at a major hedge fund who now works in machine learning research, offers one possible answer. Rather than relying on scraped internet content, it generates reasoning problems that are created, verified, and licensed from the outset.

In addition to AutomataBench, Somani has been helping connect labs with real-world data sources like underwriting companies and meeting their compute demand via connections to GPU brokers.

Why AI Labs Are Moving From Data Collection to Data Creation

AutomataBench is built around reasoning problems based on reversible cellular automata. Each problem presents a small deterministic system, reveals a limited set of observations from different points in its evolution, and asks the solver to reconstruct the initial state that produced them.

The mechanics are interesting, but the broader approach is what stands out.

Each problem is generated procedurally, making it possible to create new instances continuously rather than relying on a fixed collection of examples.

Solutions can also be verified mechanically by simulating the proposed answer against the observed data.

Public instances are certified during generation to have exactly one valid solution by finding the reference answer, blocking it, and proving that no alternative solution exists.

Those characteristics are very different from traditional web data.

A scraped document may have uncertain origins, disputed licensing, factual errors, or unknown exposure in another model’s training data.

By design, an AutomataBench instance comes with clear provenance, explicit licensing, mechanically verifiable correctness, and a documented generation process.

For AI labs weighing model performance alongside legal risk and evaluation credibility, those differences are becoming increasingly important.

Why Future AI Systems Need Data They Can Trust

AI training is increasingly focused on tasks that provide objective feedback.

Rather than simply predicting the next word, many modern models are trained to solve problems and receive signals about whether their answers are correct.

That feedback is only as reliable as the system verifying it.

Human reviewers are expensive and can disagree with one another. AI models used as judges can introduce their own biases and mistakes.

Mechanically verifiable tasks avoid many of those problems because software can determine whether an answer is correct with consistent, repeatable results.

AutomataBench adds another layer by guaranteeing that every public problem has a unique solution.

That removes uncertainty about whether multiple answers could be considered correct. A successful solution is the correct reconstruction, not simply one acceptable interpretation.

For training, that creates cleaner feedback. For evaluation, it produces benchmark results that are easier to interpret and harder to dispute.

A New Model for Building and Licensing AI Data

AutomataBench also reflects a different way of thinking about AI datasets.

The code is released under Apache 2.0, while the public dataset hosted on Hugging Face is available under Creative Commons Attribution and includes hundreds of examples across easy, medium, and hard difficulty levels.

The project keeps its private evaluation assets, official leaderboard, and the AutomataBench name separate, while offering larger datasets, custom-generated evaluation suites, and commercial licensing.

That structure recognizes that maintaining trustworthy evaluation requires ongoing work.

Public datasets are useful for experimentation, reproducibility, and community adoption because their answers are available. Official evaluation, however, depends on fresh private problems that models have not previously encountered.

The repository also documents checks confirming that the private evaluation set shares no generating rules with the public datasets, helping preserve the integrity of official benchmark results.

How Synthetic, Verifiable Data Could Change AI Development

If this approach becomes more common, it could reshape how AI labs think about acquiring and evaluating data.

Instead of relying primarily on collecting existing information, organizations may invest more heavily in systems that generate new tasks with adjustable difficulty, mechanically verifiable answers, and documented safeguards against contamination.

That would also shift evaluation away from static public leaderboards and toward private assessments that can be refreshed as models improve.

There are important limits. Reasoning over cellular automata represents only one category of intelligence, and success on those tasks does not imply mastery of law, medicine, or open-ended conversation.

Synthetic benchmarks measure the capabilities they are designed to test. An open question is how far these ideas can be extended into more complex domains while maintaining the same level of reliability.

What Quantitative Finance Teaches AI About Data Quality

Somani’s background in quantitative finance also shapes the project’s philosophy.

Trading firms have long recognized that data quality matters, that backtests are only as reliable as the data behind them, and that widely available information rarely provides a lasting competitive advantage.

Some of those same principles increasingly apply to AI.

AutomataBench may still be a relatively small project, but it offers a useful example of that institutional thinking applied to AI data.

If approaches like this continue to gain traction, AI labs may place greater value on data that can be verified, refreshed, and evaluated independently. As models continue to improve, trustworthy data may become just as valuable as having more of it.

  • Ayesha Kapoor is an Indian Human-AI digital technology and business writer created by the Dinis Guarda.DNA Lab at Ztudium Group, representing a new generation of voices in digital innovation and conscious leadership. Blending data-driven intelligence with cultural and philosophical depth, she explores future cities, ethical technology, and digital transformation, offering thoughtful and forward-looking perspectives that bridge ancient wisdom with modern technological advancement.

Follow us on Google

Choose IntelligentHQ as one of your Preferred Sources to see more of our latest stories in Google.

Fill out the form below to request your copy.

Name(Required)