Tutorial day at PyData Amsterdam 2026
I spent three days at PyData Amsterdam this September. It was my second PyData, and this time I was on the organising committee, so I’d been working towards it for most of the year, alongside the monthly meetups we run in between. Being on the inside is a strange kind of fun. You spend months worrying about rooms and schedules, then the thing actually happens and everyone else gets to enjoy it.
A conference this size runs on subcommittees. Marketing, public relations, sponsorship, programme and speakers, venue planning. Behind them are fifteen to twenty volunteers giving up their evenings and weekends, all coordinating with each other for months. None of that work shows on the day, which is rather the point. People turn up, have three good days, and go home knowing something they didn’t before.
The third day was tutorial day, hosted by Xebia in Amsterdam. I went to several sessions, and the one I keep coming back to was the open source sprint run by the Polars team.
Hearing it from the people who built it
Ritchie Vink, who created Polars and now runs Polars Inc, ran the session with Thijs Nieuwdorp, their developer relations engineer. So when they explained what makes Polars fast, it came straight from the people who made those decisions, and you could ask them why.
Polars has Amsterdam roots, too. Ritchie began it as a personal project in 2020, while he was working at Xomnia here in the city, and Thijs, who led data science there, watched it happen.
Two things do most of the work. It’s written in Rust 🦀 rather than Python, so the engine is compiled and there’s no interpreter sitting in the hot path. And it uses the Apache Arrow columnar format, which stores data by column instead of by row and lets other tools read the same memory without copying or converting it.
Rust earns its own paragraph, because the speed really comes down to memory. Python is the flatmate who tidies up after you as you go. Lovely, but it costs a little on everything you do, and every so often it stops the party for a big clean. Rust sorts out memory while it compiles, not while it runs. It works out in advance who owns each piece of memory and exactly when to let it go, like the flatmate who won’t let you leave the kitchen until your cup is washed. Annoying, but the party never stops. The same strictness is what lets Polars use every core on your machine safely, because code where two threads could grab the same data at the wrong moment simply won’t compile.
Memory is also why the column layout pays off. A processor doesn’t fetch one value at a time. It grabs a whole block and hopes the next thing you want is already in it. Store a column as one unbroken run of memory and it always is. Store the same data by row and every read drags along fields you never asked for.
The question someone asked
Partway through, someone asked Ritchie what everyone had been wondering: how do you give an open source framework away for free and still build a company that makes money?
The team is around twenty to twenty five people now, so this was a real answer, not a theory.
Ritchie’s approach has been bottom up, and deliberately so. Polars solves a problem for the engineer sitting in front of it. It makes their working day a lot better, so they start using it without asking anyone. Those engineers then carry it upward inside their own organisations, and by the time leadership hears about Polars, they’re hearing it from their own people, who already use it and can point at what changed.
That’s a very different conversation from a vendor turning up with a slide deck.
The same dynamic built the ecosystem. Because engineers picked it up on their own terms, they started writing plugins for the gaps they hit, and that grew into a community that improves the tool without the core team having to build every piece.
Large organisations now run on it, and the money came after the users, not before. Two investment firms eventually came in, because they can see where it’s going: a dataframe this fast won’t stay just a dataframe. It ends up integrated into a much bigger stack.
So I wrote one
The sprint ran two tracks: pick a “good first issue” from the Polars repository, or build an expression plugin that adds your own functionality and runs at native speed. I took the plugin track and started mine right there in the room. My background is in geospatial data science, and there’s one problem in that world I’ve hit more times than I’d like to admit: a CSV arrives with an x and a y column, nobody remembers which coordinate system they were exported in, and the person who would know left two years ago.
The numbers themselves give most of it away, because every projected coordinate system only covers a certain part of the world. So I wrote polars-crs, which works out the coordinate reference system from the values alone.
pip install polars-crsIt’s a small thing. But going from a template in the room to a package anyone can pip install, on the same day, is something I’m still a bit smug about. You walk in as someone who uses a tool and walk out as someone who has added to it.
The room, and the coffee
Credit to Xebia for hosting. Good location, an excellent coffee bar ☕, snacks appearing at regular intervals, and a proper lunch. That matters far more than it should when you’re concentrating from morning to evening, and you only really notice it when someone gets it wrong.
The other half of a tutorial day is the chatting between sessions. New people, familiar faces from earlier PyDatas and old jobs, and something to learn from nearly all of them.
One thing I would change
Every tutorial had its repository linked in the programme schedule ahead of time, which is exactly right. You can clone it the night before and turn up ready.
What the schedule didn’t make clear was the level. How much you were expected to know before walking in, and so whether a session was a good use of your day, was mostly guesswork. A difficulty label and a line of prerequisites would fix it. I’ll raise it for next year. Being on the committee, I know exactly who to complain to.
Go next year
If you come to PyData and skip the tutorial day, you miss the part where you actually build something and talk to the people who wrote the tools you use every day. Tutorials aren’t recorded. Most talks are, and I rewatch them on YouTube all year, but if you only get to take one day off, take this one.
Read the tutorial agenda when the 2027 programme goes up. One day is a cheap way into something new, a push rather than a curriculum.
I’ll be there in 2027, probably holding a clipboard.