Four weeks, 1.3 million rows, and no ways of working

I've recently spent 4 weeks going over all our core usage data that Moodle tracks — we're at roughly 1.3 million lines of data, give or take 50,000. But what's 50,000 lines between friends — in this case Claude and I.

It's been one of the most enjoyable pieces of work I've done in years; surprisingly more enjoyable than interviewing for a job last year — who'd have thought. I have fallen down so many rabbit holes that a product analyst / AI whizz would not have ventured down… but I went full steam into them and consequently clawed my way back.

I sometimes think that makes the experience even more enjoyable. So what happened?

Token Burn

At first I was running my queries in the cloud as a Claude Chat. At one point, I opened up a chat window and burned through 70% of my session tokens with one query.

Ooops.

So I gave myself a good talking to, fired up CoWork and moved everything locally. Which in turn gave us:

  • Low token burn
  • Audit trail as we had python scripts running on data with debug
  • Version control
  • Improved quality — engineers could read the scripts to make sure I/Claude wasn't doing anything "creative"

That last one turned out to be more important than I expected. People ask how you trust analysis that AI has been involved in. The answer is that someone who knows what they're looking at can open the script and check.

Version control

By version control… I use OneDrive. All I needed was a record of what was changing; not a full git repo.

Ways of working

Unlike product squads, I set up absolutely zero, nada, none ways of working. I just went straight in and started mucking around with data. All the usual ways of working fun bubbled up — ownership or lack of it, miscommunication, lack of understanding, no clear goals… the list goes on. But the real clanger was…

Definitions

We had definitions but these were loose and not fully thought out. They changed almost as much as my coffee beans. We got there in the end but this added a whole level of confusion that probably could have been avoided with a more methodical approach.

In Summary

Not one of these adventures is a new problem. Turns out (who'd have thought!) every single one has a name, and engineering has had the name for decades. Schemas. Source control. Precedence rules. Decision records. Tests.

We hit all of them in four weeks and solved each one in the order we tripped over it. Now is this bad? Maybe — but you learn from your mistakes right? And I learnt a lot.