Should AI Cook?
The framework behind my SheBuilds on Lovable hackathon build, and the question most GTM teams are skipping.
A few weeks ago, I applied for Season 3 of the SheBuilds on Lovable hackathon on a whim, fully expecting not to get in. So when I did, my idea was, to put it generously, very loose. The week before the hackathon started, I spent my nights talking to Claude, working out what I was actually going to build during the 48 hours. It truly felt like I was back working on my final year project for uni.
The awkward thing is that I’d been so sure of my original concept that I’d already published a Substack post about it, Quadrant Chart and all. But my daily conversations with Claude where we worked through the details of how to build a practical tool based on this concept forced me to admit the axes on that 2x2 were not quite hitting the mark and I needed to make a change right away. Claude even asked me: is the new version better, or just easier? Yikes!
Here’s where I landed, and oddly enough, it started with lunch.
Our office puts on lunch three days a week, and a good chunk of the conversation with colleagues ends up being about the food, then about what we’d cook ourselves at home. RecipeTin Eats’ Nagi comes up a lot, which gets us onto our own cooking habits. Yes, the pre-chopped veg costs more. Sometimes you just want the time back.
Eventually we got to the Thermomix. If you haven’t met one, it’s a smart device that cooks more or less end to end: pick a recipe, and it tells you what to put in the bowl, then chops and mixes and heats it all in the right order until a finished meal comes out. It’s quite expensive and comes with an ongoing subscription but some people would love one. Busy parents, anyone carrying a lot at home, anyone happy to trade the effort for the time back. But others, like me, probably wouldn’t enjoy it at all, because it takes the flair and experimentation out of cooking. There’s a particular feeling that shows up when you’ve done the thing yourself. Being a cooking innovation luddite wasn’t on my bingo card for this year but here we are.
Granted, there’s clearly a place for it. We’d all love the idea of a digital chef for the nights we can’t be bothered. But the machine works best when it has the exact recipe and the exact ingredients it expects. The moment you want to pull a meal together from whatever’s actually in the fridge, you’re back to relying on your own skills anyway.
That was the thing rattling around in my head when Should AI Cook? came together. The questions we kept circling at lunch, is it worth the cost, will I enjoy the result, am I losing something if I let the machine do it, are exactly the questions GTM teams are not asking before they hand work to AI. Especially in enterprise, where the data is never all there and the systems rarely talk to each other. You end up improvising with what you’ve got, relying on your own skills and adaptability anyway.
So what was wrong with my original concept?
The original 2x2 I’d published plotted a task on two axes: How AI Ready it was, and how much mastery you’d lose by handing it over. But the issue is that this format was trying to combine three concepts into a neat 2x2. The idea of “AI Readiness” is actually two separate concepts that deserve separate treatments.
Whether AI can do a task (Codifiability) and whether your organisation will actually get value from automating it (Capturability) are not the same question. A model can do something brilliantly and still be the wrong thing to automate, because it runs once a quarter, or it costs more in tokens than the output is worth, or the result goes stale before anyone uses it. Mashing both into a single “AI Ready” axis hides the decision a manager needs to make.
Splitting this into two results in a simpler framework:
Can AI actually do this step well? Also called codifiability. How consistent is the input, how bounded is the task, how repeatable is a good output. A step where everyone follows the same recipe scores high. A step where individual flair or unique knowledge changes the output every time scores low.
Will your organisation actually capture the value? This refers to capturability. How often does this run, how much input and how diverse are the inputs to be processed, how long does the output stay useful. This is the question most AI tools skip entirely, until the bills get too high and then organisations go from tokenmaxxing to tokenminimising.
But wait, what about mastery?
This was the hardest call I had to make the Friday before I started the Hackathon. When mastery was an axis, it could tell you how much judgment a task built. A single number, or a scale from low to high. But the more I considered this and tested it out with different GTM workflows and tasks, the more I realised that “you’ll lose some mastery” is not super actionable information. Lose what, exactly? And then what do I do? The mastery you build doing analysis is not the same as managing a relationship, or crafting new messaging. A tool that treats them as the same quantity doesn’t tell you anything you can act on.
So mastery stopped being a value and became a category. Not how much skill you’d lose, but what kind. For the SheBuilds Hackathon, I landed on seven types, each building a different kind of judgment:
Operational: Running a known process. Pulling account history, formatting a report, following a fixed set of steps.
Analytical: Finding the signal in a pile of inputs. Research, pattern-spotting, working out what the data is actually saying.
Generative: Making something new. Messaging, content, creative assets.
Strategic: Choosing a direction. ICP definition, positioning, deciding what to do rather than how to do it.
Relational: Reading and aligning people. Stakeholder management, negotiation, anything requiring understanding and managing people’s expectations.
Synthesis: Pulling scattered inputs into one coherent thing. Distilling a dozen calls into three themes.
Evaluative: Making a judgement call against criteria. Qualifying a deal, assessing a candidate, deciding if something is good enough.
This gives you more grain. It pulls apart kinds of work that sound identical until you look closely. For example, an analytical step and a synthesis step in a process might sound very similar on the surface but when you ask someone whether the challenge lies in finding patterns in a large set of data or summarising multiple sources into one coherent document, then you can better understand where to place AI for maximum effectiveness.
In the same way, some tasks like Operational work, might not build mastery at all. They can be ‘safely’ automated, freeing up time for your team to spend on other areas of mastery like relational or generative work.
Once someone describes their workflow and the subtasks within it, each step is tagged with an archetype which then routes to the right question about that particular step’s nuances in your workflow or company. In this way, Should AI Cook? guides your decision so the call you make is one you could actually defend to your team and still stand behind a year later.
The four quadrants
Once a step has a codifiability score and a capturability score, it lands in one of four places:
Off the Menu (low / low) — AI can’t do it well, and the value isn’t there anyway. Leave it. Not every task needs an automation decision.
Too Expensive to Cook (high / low) — AI can technically do it, but it runs too rarely, or costs too much to process.
Signature Dish (low / high) — worth doing well, but AI can’t yet. Usually the relational, strategic and evaluative work. This is what your team is for.
Crowd Favourite (high / high) — AI can do it and the value is real. The clear, uncomplicated yes.
Plot a whole workflow and you get a map: a numbered path showing where each step falls, coloured by archetype.
The Trap
So Crowd Favourite is the clean yes. AI can do it and the value to the company is real. Except sometimes the clean yes is the dangerous one.
A step can score high on both axes and be the place your people were developing their role-specific skills and experience. The economics might point towards AI automation but the archetype says: this is where your team learned to read deal signals, where your marketers learned to connect your message to pop culture trends and where someone built the instinct you can’t create a micro-learning for. The wholesale automation of these tasks doesn’t cause any immediate issues (in fact, it’s the opposite in the short term). But six months, one year down the line the team will have forgotten the foundations and will depend on AI to read their own notes and analyse the meeting they were just in to figure out the next steps.
This is the Trap. A step that lands in Crowd Favourite while carrying one of the six mastery-building archetypes is flagged by the tool. But my intention wasn’t to give leaders a concerning verdict and send them on their way. So the tool also follows up with guidance on how to keep some automation along with some desirable friction to ensure that your team retains their core skills. The steps are simple enough to work into your existing workflow: Stand up a review where someone still has to form their own view before checking it against the AI’s output. Make a practice of drafting first, comparing second, so the muscle stays warm.
So, should AI cook?
When I learn to cook new recipes, I start by following the recipe exactly. When Nagi says exactly 1 and 1/2 teaspoons of paprika, that’s what I do. I don’t trust myself to eyeball it yet, and the first few times the dish might be delicious, but it’s still Nagi’s recipe, not mine. Then one day I’m standing in front of the fridge with no recipe and half the ingredients, and I have to make something that works anyway. That moment, the riffing, and the experimenting only exists because of all the times I followed the steps first.
That’s the part I worry about us automating away. The people who’ve spent years in GTM already have their version of it: they know what a good account plan feels like, what a message needs before it ships, when a deal is about to return from the dead. And the advice everyone’s giving them now is right, feed that judgment into the AI, build the brain, get the machine to cook the way you would. The trouble is the next person. The junior who used to learn by doing the unglamorous reps now gets an average-quality output in four seconds, before the struggle that would have taught them anything. They never get to the open-fridge-door moment. They just get their food delivered, I guess.
That’s what Should AI Cook? is really for. It won’t tell you AI is good or bad, but it helps you see which parts of the work are worth keeping in human hands long enough for the next person to actually learn them. If you’re in the GTM space, run a real workflow through it, and see what comes back.
And then tell me what you’d change about it. That’s really my ulterior motive here, to get loads more testers so that I can improve on what I built during the Hackathon. There’s only so much testing I could do during the 48 hours so now I want to know which steps the tool mis-scored, which archetype felt off for your workflow, where the advice was a generic motherhood statement.
The roadmap I’m working through goes well past the first decision-support tool. I want to build coaching drills, objection handling, and competitive role-play practice that can help reps practice their skills and be more productive. Eventually I want to build a version that plugs into the CRM and workflows that reps actually live in and can deploy the right enablement for the moment. For that, I need to watch more people use the rough version and figure out where it breaks.
So this is my humble plea. Try it, break it, and tell me what you’d want it to do next. should-ai-cook.lovable.app







