The State of Agentic Data Science: From Hype to Real-World Impact
An agent can write the code, run the analysis, and explain the result convincingly. Whether that result supports a business decision is a much harder question.
On October 2, PyMC Labs, the Bayesian modeling and AI consultancy where I work, hosted a panel on the state of agentic data science. The speakers were Gaël Varoquaux from Probabl and Inria, Shipra Arora from Bain & Company, Hugo Bowne-Anderson from Vanishing Gradients, and me, Luca Fiaschi.
We discussed where agents are useful today, where they mislead us, and how the work changes when stakeholders can explore models themselves. This is my recap of the conversation.
The work moves toward defining the problem
We started by asking what data science means when agents can do more of the implementation. Gaël brought the discussion back to the specific data in front of us. A language model arrives with knowledge from its training; a data scientist needs to understand what this particular dataset can tell us.
For me, the purpose remains the same: use the scientific method to understand a problem and improve a decision or business process. Agentic data science changes the techniques and the scale at which we can work.
Shipra described the shift in a data scientist’s day. When agents take on coding and experimentation, more of the human effort goes into specifying the problem and translating the output into business action. You need to explain what success means before you delegate, then assess whether the result is useful afterward.
That is also where I see two opportunities: agents can help experts carry out analysis faster, and they can help stakeholders interact with models that previously required an expert to interpret.
Verification belongs inside the workflow
Gaël warned about the illusion of velocity. Experienced data scientists can get more done with agents. People without that expertise can also feel more productive while missing mistakes that invalidate the analysis.
His survival-analysis example makes this concrete. Suppose you are predicting how long until something happens, but some people have not yet experienced the outcome by the end of the observation period. Dropping those people can produce a beautiful accuracy number and a biased prediction. The code runs. The evaluation still fails to represent the problem.
This is why checking that an agent completed its task is insufficient. We need to check how it framed the analysis, what assumptions it made, and whether its evidence supports the conclusion.
Hugo described three layers of verification that he, Thomas Wiecki, and I teach in Master Agentic Data Science:
- Programmatic checks for properties we can test directly, such as data validity and evaluation rules.
- Agent critique to challenge another agent’s choices and investigate alternative explanations.
- Human review to judge whether the work answers the business question and whether its consequences are acceptable.
Shipra emphasized that the balance depends on the risk of the decision. An internal productivity workflow and a system producing material that reaches customers need different levels of oversight. Delegating verification is useful, but we still need to decide which checks deserve our trust.
Hugo, Thomas Wiecki, and I develop this approach in The Agentic Data Science Playbook on O’Reilly, with examples of specifying investigations, checking agent outputs, and carrying what a team learns into its next analysis.
Give agents context that survives a model upgrade
One audience question asked how workflows should evolve as models improve. Some instructions that help today’s model can become unnecessary or even restrictive for the next one.
I draw a distinction between context about the problem and procedural instructions about how to solve it.
The problem context tends to age well. What will the model be used for? Which features will actually be available at prediction time? Is there a particular customer segment whose performance matters most? A stronger model cannot infer those requirements reliably from a dataset alone.
Procedural instructions need more frequent reassessment. In our course, we use a fraud-modeling example where an agent can miss the need for a temporal train-test split. More capable models may recognize that requirement without being told. But they still need to know how the business will use the resulting model.
My recommendation in the panel was to evaluate those procedural constraints as new models arrive. Keep the business objective explicit, and check whether the prescribed method still improves the result.
Let stakeholders question the model
Gaël and I approached access to analytics from different directions. He emphasized the risk that people find patterns they want to see rather than patterns supported by the data. I agree that the risk is real.
I also think there is value in letting business users explore. In the organizations I have worked in, stakeholders sometimes found insights a data team would have missed because they understood the problem domain better. They also made mistakes. The quality of the tools and the verification built into them matters enormously.
That leads to a change in what a data scientist delivers. In marketing mix modeling, the deliverable has often been a quarterly presentation or a dashboard with a fixed simulator. With an agent around the model, stakeholders can ask how it was constructed, question assumptions, and request simulations in ordinary language.
Shipra described why that access matters in marketing: more frequent interaction with a model can help business users adopt its insights and make decisions faster. The opportunity is to deliver something people can interrogate, with checks that make the answers worth acting on.
Synthetic consumers still depend on real data
We also discussed our synthetic consumer work at PyMC Labs, including the research with Colgate-Palmolive. LLMs can simulate consumer responses, but a plausible persona is only a starting point.
In that work, we developed a method to better align synthetic rating distributions with human responses. As we worked with other data providers, we found that the information used to construct the personas mattered. Transaction histories, for example, can help ground responses about price sensitivity.
I see promise here, but generating responses does not remove the need to validate them against real behavior. The same question applies as in the rest of the panel: what evidence makes the output trustworthy for this use?
Faster analysis needs a useful destination
When Hugo asked what we could do beyond accelerating existing tasks, Gaël pointed to the backlog of experiments that teams already have data for but never find time to try. I focused on interactive data products that let stakeholders explore the reasoning behind a conclusion.
Both possibilities depend on evaluating more than speed. Cost and response time matter when hundreds of people use a system. Reliability, explainability, and avoiding analytical mistakes matter when agents build models across products and markets. The right measures follow from the decision the system is supposed to improve.
I want agents to make more questions practical to investigate, and to put the resulting models in the hands of people who can act on them. That requires keeping verification close to the work as we give agents more responsibility.
Watch the full panel recording on YouTube.