A Build-Versus-Buy Decision Sheet You Can Defend
A decision is defensible when somebody who disagrees with it can see exactly which input they disagree with. This sheet turns build versus buy into criteria, weights, scores and a dated record, so the argument moves from opinions to a specific cell.
- Applies to
- A specific automation requirement with a named run period, not a general policy
- Inputs you supply
- Criterion weights, scores for each route, cost totals from your own sheets
- Produces
- A weighted verdict, a written record of assumptions, and a review trigger
- Out of scope
- Vendor selection; this sheet decides the route, not the product
A defensible build-versus-buy decision has five parts: a stated period, weights set before scoring, a score with a written reason in every cell, cost totals from your own sheets, and a dated record naming the trigger for review. Everything else is a conversation that will be repeated in three months with different participants and no memory of the first one.
The sheet below does not tell you what to choose. It makes the choice locatable, which is the achievable goal. When somebody disagrees, they should be able to point at a cell, name the input they would change, and see the verdict move or not move. That is the difference between a decision and an argument.
Why a sheet rather than a discussion
Discussions about building are unusually prone to drift, because both sides are arguing about different things. One person is talking about capability, another about cost, a third about control, and the words used are the same. A sheet forces each concern into its own row, where its weight becomes visible and arguable.
The second reason is durability. Six months later nobody remembers which assumptions carried the decision, so nobody notices when one of them stops being true. A record with inputs printed on it is the only mechanism that catches an expired assumption, and expired assumptions are how sensible decisions become indefensible policies.
The nine criteria
These cover the dimensions that genuinely separate the two routes. Cut any that does not apply to your situation rather than scoring it zero, and add at most one or two that are specific to you. A sheet with twenty rows produces false precision.
- Time to first useful run. How long until the thing is operating with limits, logs and a stop switch, not how long until it works once.
- Total cost over the period. Both sheets, same period, same hourly rate applied to hours on both sides.
- Control over behaviour. Whether you need to change sequencing, sizing or timing in ways a hosted option does not expose.
- Operational load. Hours per month of ownership, including the ones that arrive unscheduled.
- Coverage. The venues and cases you need supported now, plus the ones you can foresee needing.
- Evidence quality. What record each route leaves that you can check independently afterwards.
- Continuity. What happens when the person who understands it is unavailable, or when a vendor changes terms.
- Disclosure. Whether your approach can be described to a third party at all.
- Reversibility. The cost and time of undoing this choice if it turns out wrong.
Choosing weights before scoring
Weights total one hundred and are set before either column is scored. This sequence is not bureaucratic fussiness; it is the entire integrity of the method. Weights chosen after the scores are visible are a conclusion wearing the costume of an input, and everyone reading the sheet later can tell.
Weight setting is also where the real disagreement lives, and it is better to have that argument explicitly. If one person weights control at thirty and another at five, no amount of careful scoring will reconcile them, and discovering that in ten minutes is cheaper than discovering it after a build.
A useful discipline is to write one sentence justifying any weight above fifteen. High weights are claims about what matters most, and a claim that cannot be defended in a sentence is usually a preference that has been promoted quietly.
The scoring matrix
Score each route from one to five per criterion, where five is better. Write a one-line reason in every cell. The reasons are what make the sheet reviewable, and a cell with a score and no reason should be treated as blank.
| Criterion | Weight | Build score | Buy score | Reason recorded in the cell |
|---|---|---|---|---|
| Time to first useful run | - | - | - | What "useful" means here, and what each route needs to reach it |
| Total cost over the period | - | - | - | The two totals and the period they cover |
| Control over behaviour | - | - | - | The specific behaviour you need to change, named |
| Operational load | - | - | - | Monthly hours from your ownership matrix |
| Coverage | - | - | - | Venues and cases required, and which are missing on each side |
| Evidence quality | - | - | - | What record survives the run and who can verify it |
| Continuity | - | - | - | Who covers absence on one side, what the terms say on the other |
| Disclosure | - | - | - | Whether anything about the approach cannot be shared |
| Reversibility | - | - | - | Cost and elapsed time to undo this choice |
A worked example with placeholder scores
Everything in this example is illustrative. The weights and scores are placeholders chosen to demonstrate the arithmetic, and they describe no real team, product or situation. Replace every figure before drawing any conclusion.
Illustrative weighted totals
Weights: time 15, cost 20, control 10, operational load 15, coverage 10, evidence 10, continuity 10, disclosure 5, reversibility 5. Total 100.
Build scores: 2, 2, 5, 2, 4, 4, 2, 5, 2. Weighted: 30 + 40 + 50 + 30 + 40 + 40 + 20 + 25 + 10 = 285.
Buy scores: 5, 4, 2, 5, 4, 3, 4, 2, 5. Weighted: 75 + 80 + 20 + 75 + 40 + 30 + 40 + 10 + 25 = 395.
On these placeholder inputs the rented route leads by 110 points out of a maximum of 500. The gap is driven by three rows: time, operational load and reversibility. Nothing about that result generalises; what generalises is that you can now see which three rows to argue about.
Test the result before trusting it. Raise the control weight from 10 to 30 and reduce cost to 10, which is what a team building a proprietary approach would legitimately do, and the two totals move much closer. A verdict that survives reasonable weight changes is robust; a verdict that flips on a five-point change was never really produced by the sheet.
Cost is a gate, not a criterion
Cost appears as a row above, which is a compromise for simplicity. In practice it works better as a gate applied before scoring: if one route costs more than the budget over the period, it is out, and the sheet decides between the remaining options on everything else.
The reason is that cost is measured while everything else is judged. A weighted score mixes an arithmetic total with nine subjective ratings and lets the ratings dilute the measurement. Gating first keeps the measured thing decisive and leaves the sheet to do what it is good at, which is comparing things that cannot be added up.
Either way, both totals come from your own sheets rather than from anywhere else. The recurring side comes from your monthly bill and your tracked hours; the rented side comes from the actual charge for the configuration you would actually use. Sizing that configuration properly is its own question, and understanding how much volume a token needs is what stops you pricing a campaign twice the size of the one you intend to run.
Running the sheet
The whole exercise takes an afternoon if the cost sheets already exist and a week if they do not. The sequence matters more than the speed.
- State the requirement and the period. One sentence for what it must do, one number for how many months. Everything downstream depends on the period.
- Set the weights. Total one hundred, agreed before any score is written, with a sentence justifying every weight above fifteen.
- Total both cost sheets. Same period, same rate, hours included on both sides.
- Score every cell with a reason. Both columns, one to five, one line of justification each.
- Multiply and total. Then stress the result by changing the two highest weights and seeing whether the verdict holds.
- Write the record. Inputs, date, verdict, review trigger, stored where the next person will actually find it.
The decision record template
The record is the deliverable. It is short by design, because a long one does not get read and an unread record cannot catch an expired assumption.
DECISION: build or buy
DATE:
REQUIREMENT: one sentence
RUN PERIOD: months
INPUTS USED
engineer day rate:
build days to first useful run:
monthly infrastructure total:
maintenance hours per month:
attention hours per month:
hourly rate applied:
rented monthly charge:
WEIGHTS: nine criteria, totalling 100
TOTALS: build weighted / buy weighted
VERDICT: chosen route, one sentence why
BLIND SPOTS: what this sheet did not evaluate
REVIEW TRIGGER: named event or date
OWNER: who watches for the trigger Two fields carry most of the value. Blind spots forces you to name what the sheet ignored, which is where the next surprise will come from. Owner turns the review trigger from an intention into somebody's responsibility, and a trigger without an owner has never once fired on time.
Presenting it to somebody who was not there
Lead with the period and the weights, not with the verdict. Anyone hearing a conclusion first spends the rest of the conversation looking for the flaw that produced it. Anyone hearing the priorities first is usually arguing about weights within a minute, which is exactly the argument worth having.
Expect the objection to be about a single cell, and treat that as success. Change the cell in front of them, recalculate, and show whether the verdict moves. A sheet that survives that treatment is more persuasive than any amount of narrative, and a sheet that collapses under it has told you something useful for free.
Where a criterion depends on a claim about a product rather than about your own team, say so explicitly and score it as unverified until you have run it. Reading the documented behaviour of the platform underneath, such as the reference material at the Solana documentation site, tells you what is protocol-level and true regardless of tooling, which is a useful way to separate a product claim from a chain fact.
The review triggers
Set them at the moment of the decision, because afterwards nobody has the context to know what would matter. Four are worth writing on almost every sheet, plus a calendar date as a backstop in case none of them fires and everybody forgets.
One trigger is external to your team entirely: the platform underneath both routes keeps moving, and it does so on a public schedule. Operational notes such as the Anza operator documentation exist because that surface changes, and a change there can quietly alter the operational load score on the build column without anybody deciding anything.
- The run period changes, in either direction, by more than a third.
- The recurring bill moves materially, including quietly through added line items.
- The person who carries maintenance becomes unavailable or announces they are leaving.
- The capability gap closes or widens, on either side, in a way that would change a score.
- A backstop date, typically two quarters out, at which the sheet is re-run regardless.
The operational side of the decision needs its own periodic look too. Whether the route you chose is still doing what you expected is a measurement question rather than a strategy one, and comparing your own records against what a professional Solana volume bot reports for the same period is a practical way to check that the assumption underneath your verdict still holds.
How these sheets get gamed
Three patterns account for nearly all of it, and all three are detectable by reading the sheet rather than by attending the meeting. Knowing them is what makes a sheet worth producing at all, because a framework that cannot be audited is just a slower opinion.
The first is weights set after scoring, which shows up as a suspiciously decisive result with weights that have no written justification. The second is criteria chosen to favour a route already selected, visible as rows that are unusually specific to one side. The third is scoring a route nobody has run, which is the most common and the easiest to fix: run a small trial and replace the guess with an observation.
The tell that a sheet was written backwards
A sheet produced honestly has at least one row where the preferred route scores badly, and the author left it there. Sheets where the chosen route wins every single criterion are not describing a decision; they are describing a preference that has been formatted. If your matrix has no losing rows, something was adjusted.
Preferring the reversible route
When the totals are close, reversibility decides. Renting first and building later is a cheap sequence: the trial costs a period of charges, produces observed inputs for every guessed cell, and leaves nothing stranded if the answer turns out to be build. The reverse sequence wastes the build and arrives at the same knowledge later.
This is not an argument for renting in general. It is an argument about order under uncertainty, and it stops applying the moment you have real numbers. Once your cost sheets are populated from observation rather than estimate, run the matrix again and follow it wherever it goes, including back to the route the first pass rejected.
The last discipline is the hardest: publish the verdict with its assumptions attached, and revisit it when the trigger fires rather than when somebody complains. A decision that is reviewed on schedule stays a decision. One that is defended past its expiry becomes an identity, and identities are considerably more expensive than either building or buying.
Questions this sheet gets asked
What is a build vs buy decision framework?
A structured way to make the comparison checkable: named criteria, weights set in advance, a score with a reason in each cell, cost totals over a stated period, and a written record of the assumptions. The purpose is not to remove judgement but to locate it, so that anyone who disagrees can point at the specific input they would change.
How many criteria should the sheet have?
Enough to cover the dimensions that actually differ and few enough that each one gets real thought. Nine is a workable number. Beyond a dozen the weights become arbitrary and the exercise starts producing precision it has not earned.
Who should set the weights?
Whoever carries the consequence, which in a small team is the person funding it and in a larger one is the person accountable for the outcome. Weights encode priorities, so setting them is the decision; scoring afterwards is mostly bookkeeping.
What if the two totals come out almost equal?
Treat that as information rather than a tie. Near-equal totals mean the criteria did not separate the routes, so the deciding factor is outside the sheet: usually reversibility, or which route you can start this week. Prefer the one that is easier to undo.
How do you score a route you have never run?
Not from imagination. Run a short trial, read what the option actually exposes, and score against what you observed. An unscored trial is cheaper than a wrong verdict, and it converts several guessed cells into observed ones.
Should the sheet be shown to the whole team?
Show the criteria and the weights to anybody affected, and be willing to defend both. A sheet used to announce a decision that was already made is worse than no sheet, because it spends credibility to add the appearance of rigour.
When does a decision record expire?
When any input it names changes materially: the period, the rate, the recurring bill, the availability of the person who would maintain it, or the capability of the rented option. Set a calendar date as a backstop, because most of those changes arrive quietly.
Filed under Build sheet by The Build Sheet Desk. Every total on this page is illustrative arithmetic built from inputs printed beside it; none of it is a quote, a benchmark or a figure taken from a real account. How the desk assembles a sheet is written out in the method note.