Building an AI-Native Account Prioritization Engine
Designing a workflow that turns GTM judgment into an operational system.
Could Judgment Actually Be Operationalized?
Part II of two. Part I asked whether GTM value was shifting from signal to judgment. This is what happened when I tried to build the system that would test it.
Part I ended on a specific, unresolved question: if judgment really is the scarce resource in GTM now — not the signal, but the decision logic sitting on top of it — what would a system built to operationalize that judgment actually look like? This is that system, built and tested against a real market, with real deadlines and no do-overs.
The Question
The question I set out to test was narrower than it sounds: could judgment — the specific, layered set of decisions a good rep makes about an account, mostly invisibly — actually be pulled out of a rep's head and rebuilt as an explicit system? Not automated in the sense of removed. Operationalized, in the sense of made repeatable, auditable, and durable enough to survive the person who currently carries it leaving.
That's what I tried to find out. Not by arguing for it. By building it, against a real company, and watching where it held and where it didn't.
I want to be specific about how this was built, because the constraints are part of why I trust the result.
This wasn't built from a blank page with unlimited time. It was built during Anthropic's AlphaForge GTM Engineering Bootcamp, across three sequential, independently graded stages, each with its own deadline and its own reflection due before I could move to the next one. I didn't get to redesign the market after the fact if an earlier decision turned out to be wrong — whatever I shipped at each stage is what the next stage had to build on.
The market itself was real, not a sandboxed exercise: ZoomInfo's own addressable market, the same mid-market and growth-stage B2B SaaS and services companies I analyzed in Part I. The data was real — sourced and enriched against actual companies, not a synthetic demo set. And the problem was the actual GTM problem, not a simplified version of it: take a broad market, decide who belongs in it, decide how much attention each account deserves, decide how it should be worked, and then actually try to produce a real interaction with a real person.
None of that guarantees the system is good. But it means the failures below are real failures, not hypothetical ones I get to wave away.
The Framework
The operating model I built has five layers, in order: Signal → Qualification → Prioritization → Segmentation → Action.
Each layer answers a different question. Each question, in a normal revenue org, gets answered — but inconsistently, invisibly, and inside individual people rather than inside anything the org can inspect or improve. The bet underneath this whole project is that if you can make each of those five questions explicit, you've operationalized the judgment. If you can't — if the system still needs a person to quietly fill in the gaps — you haven't.

Want to see the system in action: Watch a 5-minute walkthrough of the working Clay workflow built during this project.
Signal. Before anything can be qualified or scored, you need actual evidence about a company, not an assumption. I pulled two kinds: firmographic and behavioral evidence — employee count, whether a CRM was in place, whether a sales-engagement or intelligence tool sat on top of it — sourced from LinkedIn and BuiltWith. This is the raw material everything downstream depends on. In a normal org, this step happens ad hoc and per-rep: someone clicks through LinkedIn and a tech-stack tool right before — or instead of — reaching out, and nothing about what they found gets captured anywhere. Later in the build, a second kind of signal shows up again, at the point of action: a live, per-account "why now" observation generated at outreach time, not baked into the upfront score. I'll get to that.
Qualification. The first real decision: does this company belong in the market at all? I gated on geography, industry, B2B status, employee floor, and competitor exclusion — run once, upstream, before an account was scored on anything else. This is the decision a rep almost never gets to make explicitly. Most reps inherit a list someone else built and don't know, and can't ask, why a given company is or isn't on it. Making this a hard, visible gate — instead of an assumption baked silently into whoever built the list — was the first piece of judgment I tried to pull out into the open.
Prioritization. Given that a company belongs, how strong is the fit? I scored qualified accounts on two proxies: company scale (employee count) and revenue-tooling maturity (whether the company had already invested in the kind of stack that outbound depends on). Out of 1,441 sourced and qualified companies, the accounts that scored highest formed the prioritized list downstream stages would actually work. This is the decision senior reps make well by instinct and new reps don't make at all — "which of my qualified accounts actually deserves my time first" is usually gut feel, not anything the org could audit or teach.
Segmentation. Given the fit is strong, how should this account actually be worked? This is where I think the project earned its most durable lesson, so I'll say the finding plainly here and come back to it in "What I Learned": a segment only counts if it changes who gets contacted or how. I routed the prioritized market — 482 accounts once the market had been fully qualified and scored — into three segments by employee count, on the theory that headcount is a proxy for a real underlying thing: how many stakeholders are actually involved in a buying decision at that size. Roughly 40% of accounts fell into a segment where one sales leader is the whole buying committee. Another third landed where a sales leader and a RevOps counterpart both have to be involved. The rest required an executive buyer and a RevOps leader together. Each segment got its own stakeholder map and its own outreach playbook — not just a different score, a genuinely different motion. This is the decision that normally lives entirely inside a senior rep's informal playbook, and is exactly what disappears when that rep leaves: nobody documents "I work small accounts differently than big ones, and here's specifically how," they just do it.
Action. Judgment isn't real until it produces an actual interaction with an actual person. I picked the simplest of the three segments — the single-stakeholder one — specifically to isolate one variable at a time, and selected the top-scoring accounts within it. For each one, I generated a live "why now" observation from recent company activity, drafted a short, specific outreach message tied to that observation, and routed every single one through a human review step before anything went out — nothing sent itself. This is the decision reps make under the most time pressure and get wrong the most often: is this specific enough to actually earn a response, or am I about to send something generic because prep takes too long to do properly. Every send was tracked — who, what signal, what was sent, whether a connection request went out. What I did not build, and want to be direct about, is a way to capture or classify what happened after. The outreach went out. Whether it actually worked is a question the system, as built, can't yet answer for itself.
What Broke
This is the section I trust the most, because it's the one place a builder can't fake it.
The first break was structural, and it happened early. My first pass at qualification and prioritization used one combined score — geography, industry, competitor status, scale, and tooling maturity all sitting in the same calculation. It worked cleanly on a small test sample. It broke down completely once I ran it at real scale, against the full market. Not because the logic was wrong — because of cost. Checks that are free to run twenty times are not free to run fourteen hundred times, and the combined model was re-running qualification logic on every row, every time, when it only ever needed to run once, upstream, as a gate. I had to tear the model apart and rebuild it as two separate layers before it would hold. That single failure is the reason "Signal → Qualification → Prioritization" exists as three distinct steps in this write-up instead of one.
The second break was quieter, and I only fully understood it in hindsight. Between finishing the qualification-and-scoring stage and starting segmentation, the qualified market grew — not by much, but enough that the two stages report slightly different totals for how many accounts made the cut. The most honest explanation I have is that I kept improving the scoring underneath the surface between stages, prioritizing getting the foundation right over freezing the number for a clean write-up. I'd rather show that seam than pretend the number never moved.
The third break was in my own assumptions about what segmentation was actually for. Going in, I thought the interesting output of segmentation was the routing logic itself — the rule that sorts accounts into buckets. It wasn't. The real test turned out to be the playbooks: could I say, specifically, how a rep would work each segment differently? Several segmentation ideas that looked reasonable on paper — splitting by industry, by region, by how much tooling a company already had — didn't survive that test, because none of them actually changed who got contacted or what got said. I cut all of them. I also noticed, honestly, that my instinct under no time pressure is to keep refining a system past the point it's already good enough — this project is where I caught myself doing that and made myself stop.
The fourth break was the one that mattered most, and it showed up last. I assumed, going into the execution stage, that the hard part would be generating a message specific enough to earn a response. It wasn't — that part came together faster than I expected. The actual hard part, which I didn't see until after messages were already going out, was that I had built a system that could produce outreach but had no way to learn from what came back. I hadn't defined what a good outcome even looked like, hadn't built anywhere to record it, and hadn't decided how — or whether — a real response should change anything upstream. I'd designed judgment for the moment of sending. I hadn't designed judgment for the moment of finding out whether sending was right.
There's a smaller, more personal break worth including here too. My first drafts of outreach copy got more "sophisticated" — more strategic-sounding, more industry-specific — the longer I worked on them, and the more sophisticated they got, the less honest they felt. I realized I was writing lines I wouldn't actually know how to follow up on if someone replied. I threw those drafts out and set a much blunter rule for myself: if I wouldn't be comfortable continuing the conversation after a reply, I shouldn't send the message. That's not a technical constraint. It's the adoption problem showing up in its smallest, most personal form — a system only works if the person running it actually trusts what it's about to say on their behalf.
What I Learned
Strip away everything specific to this particular build, and a few things are left that I think are true about GTM operating models generally, not just about this one.
Qualification, prioritization, and segmentation are three different questions, and collapsing them into one score doesn't just look untidy — it has a real, compounding cost the moment you're operating at real scale. I learned this by having it break, not by reasoning my way to it.
A segment is only real if it changes the motion. If the stakeholder, the channel, or the message wouldn't actually be different, it's not a segment — it's a filter dressed up as one. That test is simple enough to apply to almost anything a GTM org calls a "segment," and I'd guess a lot of what gets called segmentation in practice would fail it.
The hardest part of this kind of system isn't generating information. It's producing something specific enough that a real person responds to it. Research, scoring, and stakeholder mapping can all be made deterministic and automated with real confidence. Deciding what's actually worth saying to a specific person, right now, is where the system runs out and judgment has to take over.
And the biggest one: judgment doesn't only happen at the point of action. It also has to happen in deciding whether the action worked. A system that can produce good-looking outreach but can't tell a useful response from a useless one hasn't actually operationalized judgment — it's operationalized guessing, with better production values. I built the first kind of judgment into this system. I did not build the second kind. That gap is the most honest thing I can say about the current state of the project.
Return to the Original Hypothesis
Part I's hypothesis was that GTM value is moving from collecting signal to operationalizing judgment about it — and that the winning systems would be the ones that turn "what matters, how much, and what happens next" into something repeatable, auditable, and durable enough to survive a rep leaving.
This build supports that hypothesis, at least partially, and I want to be precise about how far. I was able to take a real market, make its qualification logic explicit and auditable instead of implicit and untraceable. I was able to make prioritization a defensible score instead of a gut call. I was able to turn "how should this account be worked" into a documented, testable rule instead of a senior rep's private playbook. And I was able to get all the way to a real, human-reviewed action — not just a plan for one. That's real evidence that judgment, at least at the level of qualification, prioritization, segmentation, and the decision to act, can be pulled out of a person's head and rebuilt as a system other people could run, inspect, and improve.
What I don't have evidence for yet is the harder claim: that the system can learn. I don't know whether the outreach worked, because I never built the layer that would tell me. And the deeper problem that gap exposed isn't really technical — generating structured response categories or a tracking table isn't hard. The actual problem is organizational: deciding, as a matter of process, what counts as a meaningful outcome, who's responsible for capturing it, and how often the upstream logic actually gets revisited because of it. That's not a Clay limitation. It's the same kind of judgment problem this entire project was trying to solve, showing up one layer higher than I expected it to.
So here's where I've landed. This experiment made me more confident that judgment can be operationalized — the qualification-through-action chain is real, and it held up under a real market and real constraints. But it also showed me how much of that work is organizational, not technical. Building the system that decides was the easier half. Building the system that learns whether it decided well is the half I haven't done yet.