It grades a plan, PRD, or roadmap on priority-definition quality — specifically: actionability, priority clarity, risk coverage, measurability, and dependency awareness — and returns three artifacts: a scorecard, a line-cited critique, and a tightened rewrite.
This is intentionally not a "general code/document assistant." It is a specialist that does one thing: tell you whether a plan is a real plan, and where it is not.
Two observations.
One. The post that inspired this hiring round names "priority definition ability" as the single essential quality. The post argues — correctly, in my view — that AI has collapsed the cost of executing a plan to nearly zero, and so the differentiation has shifted entirely to which plan, in what order, and why. The scarce resource is no longer effort. It is clarity.
Two. Default Cursor's Claude is excellent at writing a plan in one shot. But that same one-shot model is a poor judge of plan quality, because:
- It has no enforced rubric and slips into vibes.
- It hallucinates items that are not in the document.
- It cannot show its work; you have to take its score on faith.
- It produces different scores on different runs of the same input.
A plan-writing tool and a plan-judging tool are different products, and the company's stated #1 quality is the one that needs the judge.
Three reasons, in order.
-
It mirrors the rubric the company themselves published. They are hiring for priority-definition ability. An agent that grades priority- definition quality is the most direct demonstration possible of that exact skill.
-
It is defensible. I can ground-truth it with a calibration set of ~30 plans hand-rated as good/medium/bad, and I can show, with numbers, where my agent agrees with humans, where default Cursor's Claude does not, and how much variance each one has across runs. Most agent submissions cannot make a defensible quantitative claim. This one can.
-
It is buildable in 7 days. The scope is bounded — five rubric dimensions, four pipeline stages, ~30 calibration plans. There is no integration sprawl and no infrastructure dependency. That matters because the JD says "earlier submission = positive signal."
- I do not try to write plans from scratch. That is a different product.
- I do not try to score on dimensions outside the rubric (e.g., "is this technically correct?", "is this strategically sound?"). Those require domain context this agent does not have.
- I do not try to be a multi-language agent in v1. English plans only.
These are explicit non-goals; they keep the rubric tight and the benchmark honest.
An APO is a priority-orchestration role. The work product of an APO is
a sequence of decisions about what to do next and why. priorityjudge
is an artifact-level critic of that exact work product — it reads what an
APO would produce and tells them, with citations, where it falls short of
the bar.
If I were to use this agent in my own APO workflow:
- Before any cross-team plan is shared, run it through
priorityjudgeand accept the tightening pass. The marginal cost is < $0.10 and ~30s. - Use the score distribution across a portfolio of plans as a leading indicator of execution risk. Plans clustered below 5000 are the ones that will slip; that is where attention goes first.
- Use the rubric itself as a lightweight team norm: "we don't ship plans scoring under 6000 on dependency awareness" — a concrete, falsifiable bar that everyone can apply themselves.
The agent is a tool. The norm it enables is the priority-definition discipline.