PyData Amsterdam 2026: agents in production, and who pays for them

pydata
agents
llm
machine learning
My takeaways from the two conference days: data modelling’s third comeback, agents as user interfaces, the token bill, and big models doing the heavy lifting for small ones.
Published

September 14, 2026

For most of this year I worked with around fifteen other committee members to organise PyData Amsterdam. We split the work into subcommittees for sponsorship, the programme, proposals, volunteers, marketing, the website, ticketing and more. I did it with a lot of joy, and learned a great deal along the way. Organising is my way of giving back to the open source community, since I use a lot of open source to build my business. And a conference like this is where younger data scientists and engineers get to see where the field is heading, and which skills will matter next.

Three hats

At the conference itself, on 10 and 11 September at the NDSM Loods, I wore three hats. 🎩

As an organiser, my job was making sure everything ran smoothly.

As a hands on co-founder, I came to learn. At my startup I build the data and LLM infrastructure, so I wanted to know which practices work, which technologies hold up and which don’t, and how companies are building AI infrastructure at a cost that still delivers real value to the business. Learning that from other people’s production incidents is a lot cheaper than having your own.

And as a co-founder building a team, I went looking for talks on how teams should collaborate to perform at their best, and how that collaboration creates satisfaction for the people in it. This year there weren’t many on team collaboration and ways of working, while last year’s edition had some very nice ones.

I watched how the talent market is moving too: which skills to look for when I hire, and how the people at the top of their field keep developing themselves. With the industry moving this fast, the shape of a strong candidate keeps changing, and I’d rather notice that in the hallway than halfway through a hiring round.

Across the two days, three trends kept coming up: evaluating LLMs properly, getting data ready for agents to use, and connecting LLM engineering back to classic machine learning. Agents are in production now, and the hard questions have moved to evaluation, data and cost.

The keynotes

The opening keynote was one of my two highlights. Maarten Grootendorst and Jay Alammar, who wrote Hands-On Large Language Models together, talked about the move from LLMs to agents: from words to actions. It made a real impression on me.

The second day opened with my other highlight, Christophe Blefari, co-founder of nao Labs, on the future of analytics.

Same job, six eras

Christophe walked through six eras of the data stack, from the enterprise data warehouse of the 1980s to the AI ready data stack of today. His verdict: it’s the same job in all six. We kept moving data around, stakeholders still ask the same questions, and SQL won.

1980sEnterprisedata warehouse2006Big data2012Cloud datawarehouse2019Lakehouse2019Embeddedanalytics2025AI readydata stackThe same job in every era: move data around, answer the same questions, in SQL
Six eras of tools, one job: moving data around and answering the same questions, in SQL. After Christophe Blefari’s keynote.

Then came a row of three tombstones for data modelling, declared dead in 2010 (“Schema on read. Just dump it in the lake.”), in 2018 (“Storage is cheap. One Big Table.”) and in 2023 (“The LLM will figure it out.”). The first ended in data swamps. The second ended in layered dbt models, which is data modelling under another name. And the third is being disproved right now: “Agents fail on raw tables. Context, documentation and modeling are back as the product.”

R.I.P.Data modelling2010“Schema on read.Just dump it in the lake.”Data swampsR.I.P.Data modelling2018“Storage is cheap.One Big Table.”Layered dbt modelsR.I.P.Data modelling2023“The LLM willfigure it out.”Agents fail on raw tables.Modelling is back.
Data modelling, declared dead three times. After Christophe Blefari’s keynote.

That fits the trend of getting data ready for agents. An agent can only answer questions about your data if the data is sound and the business context is explicit, and that context usually lives in people’s heads. Writing business context down well enough that an agent stops guessing is still human work.

Two agents on the stack

Christophe’s picture of the stack still has spreadsheets among the sources (“yes, still”, the slide admits), and two agents. One is the analytics agent, next to the dashboards, the notebooks and a human with a SQL console. The other is a platform agent that watches the stack itself: it reads failed runs, costs and schema drift, fixes the code and opens a pull request, while the data team reviews and owns the meaning. The whole stack is code, in one repository.

SOURCESPostgresSalesforceStripeEventsSpreadsheets (yes, still)ingestStoragean overgrown PostgresserveUSED BYDashboardsSemantic layerNotebooks and MLSQL by handAnalytics agentPlatform agentwatches, fixes, opens PRswatches: failed runs,costs, schema driftfixes codeOne repositorythe whole stack as codeopens a PRData teamreviews, owns meaning
One stack, two agents: one answers questions, one maintains the code. After Christophe Blefari’s keynote.

This is close to how we build at my startup. We work in Lego bricks, with small reusable blocks and pipelines that refresh one slice at a time, and we’re moving towards a software factory where agents take on more of the engineering. A platform agent that reads a failed run and opens a pull request for a human to review is a software factory for the data stack. And the tombstones are a good reminder of why the blocks matter: the model of your data is the part that survives every era.

He presented from HTML slides, like a lot of the speakers I saw, and showed where that leads: a live transcript of his talk ran in the corner of every slide and into a small database he could query with an LLM on stage. One slide was nothing but open questions. Do we need a semantic layer? Should I keep a BI tool? How much does it cost? What’s the future for people who like to code by hand? Between them, the two keynotes gave me more ideas about the future of analytics than I have free evenings to try.

Agents as user interfaces

TomTom’s talk showed what careful agent engineering looks like. Their route planning agent uses a small, fast model to talk to the user, which cuts the wait people notice, while a larger model decides which API tools to call. Tool use is prescribed explicitly instead of left to the model, and benchmark tests check whether the right tools and data were used. To go faster still, they sometimes run two models in parallel and show whichever answers first, accepting double the token cost as a deliberate trade off.

One idea stuck with me. Treat a conversational agent as a user interface. People judge it by whether they feel helped, so speed and tone matter as much as correctness. That pulls AI engineering a lot closer to software development and UX design.

The token bill

This is the part a co-founder reads twice. A speaker from Nebius, which hosts open weight models as a cloud service, argued that commercial model subscriptions are currently sold below cost. He expects that to change, and prices to rise.

His answer, unsurprisingly, is Nebius’s own Token Factory: you pay them to run an open weight model for you, often one of the strong Chinese models such as Kimi. According to him, his customers already come out much cheaper than with comparable commercial models. Self hosting is the other route, but that means a serious GPU investment plus people who can keep it running.

Take a vendor’s pitch for its own product with the usual pinch of salt. The economics are real, though. At GeoBirds we took the other route: we’ve mastered deploying our own token factory at scale, very cost effectively, for fine tuning and serving very large models.

The bridge back to classic machine learning

This was the trend I liked most, and two talks on the first day showed it from different sides.

In Distilling LLMs into Classical ML for 5,000+ Classes, Oz Mendelsohn of Swap had text to sort into more than 5,000 classes, each with a formal definition. Labelling that by hand is out of the question. So an LLM agent does the labelling offline, like a domain expert would, and gradient boosted trees learn from its labels and serve the predictions in real time, at a fraction of the cost and latency. New inputs from production get labelled and fed back in, with people auditing the results, so the accuracy keeps climbing.

Kai Jeggle of Dexter Energy took the same idea to weather in Embed First, Predict Later. The encoder of Microsoft Research’s Aurora weather model turns raw weather data into embeddings that capture things like storms and cloud systems. Small models then map those embeddings to wind, solar or demand forecasts, and one set of embeddings serves all of them.

It’s a very deterministic way of using a big model. Let the big model do the heavy lifting offline, and a small model serve the predictions. By the time anyone asks a question, the work is done, so you’re running a classic model at high speed and low cost, and getting the same answer every time.

Next to all those agents that won’t give the same answer twice, it was refreshing. This one is going on my list of things to try.

Pentest your own chatbot

One of the most popular things on the expo floor was a prompt injection challenge: five levels of talking a chatbot out of a secret code. Each level added a defence, from input filtering that blocked suspicious messages to output filtering that caught the code on the way out. The last levels took tricks like asking the bot to spell the code with spaces between the characters.

Fun game, serious lesson. Anything with a chat box is an attack surface. Defences that look solid fall to a creative enough question.

At the Polars stand

Polars had a stand showing their new query profiler, which you attach to your queries. When asked why it reports back to their servers, the team was refreshingly honest: they want to learn what people actually use Polars for, so they know whether to optimise joins or group bys. For the record, it sends query plans, not your data.

Credit where it’s due

The chairs put a lot of their own energy, and their networks, into making it happen, and the volunteers gave up a lot of their free time. The sponsors were key to its success. Without them there’s no conference, and this one matters for the whole data community in the Netherlands.

The venue was very alternative: the NDSM Loods, a former shipbuilding hall 🚢 that’s now an art space, with a lot of character.

Three tips to end with

  1. Get your data agent ready before you build the agent. It can only be as good as the context you give it.
  2. When you need the same answer twice, use the LLM to label your data, not to serve your predictions.
  3. If you manage a team, send them next year. They’ll come back with fresh ideas, and you get the credit for sending them.