Three years ago, building your own tooling meant a six month software project with a spec, a budget, and at least one engineer you could not spare. That is what made buying the conservative choice. That reason has expired, and the decision frame most organizations still use was built on it.
Buy versus build is also a door short. There are three, and naming the third one changes how the other two should be evaluated.
Buy commercial off the shelf. You are purchasing a vendor’s opinion about how your workflow should go, and along with it their roadmap, their model choices, and their subprocessor chain.
Compose. You assemble the workflow yourself on top of a general model. In practice this looks like a set of prompts or skills, a document store, a retrieval step, and a thin interface for review. You are purchasing speed and control, and taking on the maintenance.
Commission. You pay someone to build it for you. You are purchasing build capacity, and inheriting their verification discipline or their lack of it.
The middle door repriced the other two. A working version of a document review workflow that would have been a quarter of engineering in 2023 is now a few days of careful configuration. The cost did not disappear. It moved out of construction and into maintenance and verification, which are the two line items organizations are least practiced at budgeting.
A Decision You Make Repeatedly
One clarification before going further, because it changes how the rest of this reads. This is a decision you make per workflow, not per organization.
Intake summarization, BAA review, policy gap analysis, and vendor questionnaire response are four different problems with four different consequence profiles, and there is no reason they should all go through the same door. An organization running a bought tool, a composed one, and a commissioned one at the same time is not being inconsistent. It is being accurate about four things that are not alike.
What that does mean is that the sorting happens before the shopping. Two questions about each workflow do most of the work: how close does it sit to what makes your organization good at what it does, and what does a wrong output cost you?
Low differentiation with high consequence is where buying earns its keep, because you want a vendor whose other customers are finding the failure modes before you do. High on both is where commissioning, or composing with real discipline, pays for itself. High differentiation with low consequence is the natural home for composed tooling, and it is where most organizations should start. Low on both is usually a workflow to leave alone a while longer.
Managing the mix that results, the inventory, the ownership, the question of whether one verification standard applies across all of it, is its own problem and a larger one than it first appears. What follows here is how to choose well for a single workflow, which has to come first.
What Each Door Is Good For
Buying off the shelf is right when your problem is genuinely the same as everyone else’s, the vendor’s defaults describe a process you would have designed anyway, and the compliance surface is documented rather than promised. That surface has a specific shape: an executed BAA, a current subprocessor list, stated data residency, a retention and deletion policy you can read, and an explicit contractual answer on whether your data trains their models. A vendor who can hand you all five without a call is a vendor who has been asked before.
What you are accepting is that their roadmap becomes your roadmap. Over a few years, their opinion about your workflow tends to become your workflow, because the path of least resistance is to do the thing the tool does well.
Composing is right when the workflow sits close to what makes your organization good at what it does, the volume does not justify a vendor contract, and you have someone internally who can look at an output and tell whether it is correct. That last condition carries the most weight, and it is the one most often assumed rather than checked.
What you are accepting is that you own the maintenance, there is nobody to escalate to at two in the morning, and “someone internally” has to keep existing. Composed tooling has a specific failure pattern: it works beautifully until the person who built it changes roles, and then nobody can say why a particular instruction is in there or what breaks if it comes out.
Commissioning is right when the workflow is genuinely yours, it needs to run without you supervising every output, and you can specify the result well enough to hold someone to it. A good commissioned build is faster than composing and more yours than buying, which is a real combination when the conditions hold.
What you are accepting is the risk of buying a demonstration rather than a system, and a dependency you should structure deliberately rather than discover later.
The Gate That Sits In Front Of All Three
In most businesses, a wrong output from an AI tool is a bad experience. A user gets an odd recommendation, notices, and moves on. In regulated healthcare work, a wrong output is a finding. It is a control that was documented as covered and was not. It is a gap analysis that recorded a policy as sufficient when it was silent. It is a BAA that passed review with no subcontractor provision in it.
That difference changes which question decides the door. The deciding question sits upstream of price and upstream of the feature grid: can you write down what a correct output looks like, and check a system against it?
If you can, all three doors work, and you are choosing on economics, control, and how much of your own attention you want to spend. If you cannot, commissioning is the worst of the three, because you will get something that performs beautifully on the documents it was demonstrated with and no way to detect the month it stopped being right. Composing at least keeps you close enough to the outputs to notice. Buying at least gives you a vendor whose other customers are hitting the same problems.
Building The Graded Set
The practical form of that gate is a set of documents you have already worked by hand, where you know the right answer and can explain it.
Ten to twenty is enough to be useful. What matters is composition, not volume.
Capture the reasoning, not only the conclusion. “This BAA is deficient” is not a gradable answer. “This BAA is deficient because it omits the subcontractor flow-down and it caps breach notification at a period longer than the covered entity’s own obligation” is. A system that reaches the right conclusion through the wrong reading will fail on the next document, and only the reasoning tells you that in advance.
Weight the set toward the hard cases. Clean documents are not where tooling fails. It fails on the ambiguous ones: the policy that addresses a requirement obliquely, the contract that uses non-standard terminology for a standard provision, the document that is technically compliant and practically useless. If you have ever had two qualified reviewers disagree about a document, that document belongs in the set. Disagreement cases are the highest-value entries you have, because they are the ones where a plausible wrong answer is available.
The set is reusable across all three doors. It is your bake-off when you are comparing vendors. It is your acceptance criteria when you commission. It is your annual re-check when you compose. Built once, it answers the same question in every direction.
Evaluating A Builder
If you commission, five questions sort candidates faster than any portfolio review.
How will we know this still works in six months? A good answer describes a test set drawn from your documents, run on a schedule, against a threshold, with a named owner. A weak answer is “we’ll monitor it,” which means nobody will.
Do they ask what happens when it is wrong before they ask what features you want? The sequence tells you whether they have built anything that carried consequences. Builders who have shipped into regulated environments start with the failure surface because that is what determined their design. Builders who have not start with the feature list.
Can you read the logic? Prompts, rules, retrieval configuration, model selection, and the reasoning behind each. If any of it is withheld as proprietary, you have purchased a vendor dependency in a custom-build costume, at custom-build prices. This is worth settling in the contract rather than in a conversation.
Do they have domain judgment, or only AI skill? Someone who cannot distinguish a good output from a bad one in your field cannot build a system that reliably produces good ones. That arrangement can still work if you are supplying the judgment, but you should know going in that you are, and staff the review time accordingly, because it will be more than you expect.
What does month 13 cost? Get a number or get maintenance terms before signing. A build with no answer to this is a build that was priced as a project and will be experienced as a subscription.
The Maintenance Math
Construction is now the cheap part, which means the budget line most organizations fail to create is the one that keeps the system honest.
Four things change underneath a working tool. Models get deprecated and replaced, and the replacement is not always better at your specific task. Documents drift in format as the organizations producing them update their templates. Regulation moves, and a tool that encodes a requirement encodes it as of a date. The Security Rule overhaul is the working example: the proposed rule has been out since January 2025, the final rule is now targeted for 2027, and anything built against today’s text has a known expiration on it. And people leave, taking the undocumented reasoning with them.
The mitigation for all four is the same and it is unglamorous: re-run the graded set on a cadence, compare against the last run, and treat a drop as an incident rather than a curiosity. Quarterly is defensible for most workflows. After any model change is mandatory.
What You Should Own Regardless Of Door
Four things, whichever way you go.
The graded set, in a format that is yours and portable.
Readable configuration. If you cannot see the instructions the system is operating under, you cannot assess it, and neither can anyone reviewing you.
Exportable outputs in a format that outlives the tool. Findings that only exist inside a vendor’s interface are findings you lose when the contract ends.
A written, dated rationale for why you chose the door you chose. This is the artifact that answers a regulator or an enterprise reviewer asking how you decided to let a system touch protected health information. It takes an hour to write at decision time and cannot be reconstructed convincingly afterward.
Applying It
For health tech companies, this decision sits upstream of your next enterprise security review. Internal tooling with no verification story becomes a question you cannot answer in a vendor questionnaire, and the questionnaire will ask. Compose freely, but compose with the graded set from the beginning rather than adding it when a health system’s security team asks how you validate outputs.
For provider organizations, this decision usually arrives as a purchase rather than a build, which turns the five builder questions into procurement questions. They belong in the evaluation before the demo, not after it. A vendor who cannot describe how you would detect degradation is telling you something useful about how they think about their own product.
The order of operations is the same in both contexts. The graded set is the first deliverable of this project, not the last, and it is the deliverable you keep no matter which door you walk through.
Already running tooling and not sure what is out there? That is the other half of this problem, and it is a different job. AI capability enters five ways and only one of them is a purchase, so finding what is already running means probing the four paths that appear in no vendor list and no spend export. The AI Tooling Inventory skill runs that pass, and the manual workflow it automates is published alongside it.
If you are working through this decision, I would like to talk. We will go through what you are actually trying to automate and sort it into the right door.
I work on all three sides of this. I run compliance engagements, I built a platform for the mechanical parts of the work, and I help organizations put together tooling they run and maintain themselves. Which of those fits your situation is the outcome of that conversation rather than the premise of it, and I will tell you plainly when the right answer is the smallest one on the list.