AI Should Assist the Workflow, Not Control Every Product Decision
The most reliable AI products give the model a defined job inside a workflow and keep the consequential decisions with software rules and people. How we draw that line, with lessons from BossBoard.
There is a moment in almost every AI project where someone says: why not just let the AI decide? It has all the information. It is faster than a person. It never gets tired. Why build rules and review steps when the model can simply handle it?
I understand the appeal. I have felt it. And I have come to believe it is the single most expensive instinct in AI product design. Not because models are not capable, but because a product is not a collection of decisions. It is a workflow with state, responsibility and consequences, and most of those consequences do not belong to a probabilistic system.
This article is about where we draw the line at Technobita, why we draw it there, and what we learned about it while building BossBoard, a product where getting this wrong would have meant building something harmful.
Two kinds of decisions in every product
When I look at any workflow, I try to sort its decisions into two groups.
The first group is interpretation. What does this document say? Which category does this request belong to? What would a clearer version of this sentence look like? What are the key points in this conversation? These decisions are about understanding and transforming information. They are probabilistic by nature, a person would also make them with some uncertainty, and the cost of an occasional mistake is usually low because the output is still going to be looked at or used in a controlled way.
The second group is consequence. Who is allowed to do this? Has this work been approved? Is this customer's account active? Should this be published? Should this person's reliability score change? These decisions change state that other things depend on. They have to be explainable, repeatable and, in many cases, attributable to a person. A mistake here is not a slightly worse paragraph. It is a wrong permission, a wrong record, a wrong judgment about a human being.
The design rule that falls out of this is simple to say. AI can own interpretation. Software rules and people own consequence. The model can interpret, extract, generate, classify, summarise and suggest. The application controls identity, permissions, ownership, workflow state, publication and business rules.
AI may interpret, extract, generate, classify, summarise and suggest. Software controls identity, permissions, ownership, workflow, publication and business rules.
The BossBoard test
BossBoard is a product we built for small business owners who want to know who completes work properly, not just who marks tasks done. It connects positions, SOPs, tasks, proof, review and reliability into one accountability loop.
It would have been very easy, and very tempting, to make AI the judge. Let the model look at the proof a staff member attached, compare it to the SOP, and decide whether the work was done well. Let the model calculate reliability. Let the model decide who is dependable.
We did not do that, and the reason is the core of this whole argument. Reliability in BossBoard is a judgment about a person that affects how an owner sees them. That is about as consequential as a product decision gets. If a model made that judgment, nobody could fully explain why one person scored lower than another. The owner could not defend it. The staff member could not contest it. And the product would quietly become the thing it promised not to be: a surveillance and scoring machine with no accountability of its own.
So the review is done by a person, against a defined fair scoring model, and the reliability signal is calculated deterministically from reviewed outcomes. Every number can be traced back to specific reviews by specific people. That is not a limitation we accepted reluctantly. It is the product.
Where does AI fit? Upstream. Helping an owner write a clearer SOP. Flagging that a task description is ambiguous before it is assigned. Summarising a pattern in corrections so the owner can see that one procedure keeps causing problems. All of that is interpretation. All of it makes the human decisions better without replacing them. Our principle for the product is four words long: AI assists, humans decide.
Deterministic core, probabilistic edge
The architecture that follows from this is what I think of as a deterministic core with a probabilistic edge.
The core is the state machine of the product. The records, the transitions between states, the rules about who can do what, the calculations that other things depend on. This part is ordinary software. It is tested the ordinary way. Given the same inputs, it produces the same outputs, every time, and you can read the code to know why.
The edge is where the model lives. It takes something from the core, interprets or transforms it, and hands a result back. The core then validates that result before accepting it. If the result is malformed, or fails a check, or the model simply did not respond, the core already knows what to do, because the edge was never allowed to change state directly.
The practical benefit is that the intelligence can fail without the product losing the workflow. When the model times out, the task is still in the right state. When the output is nonsense, it is marked invalid rather than saved. When you change models next year, the core does not care.
The practical cost is that you have to design the boundary explicitly. You have to decide what the model is given, what it must return, and what the application does with it. That is more work than letting the model do whatever it likes. It is also the difference between a product and a demo.
Signs the AI has too much control
When we review an AI system, our own or someone else's, these are the warning signs that the line has been drawn in the wrong place.
- A model output is written directly to a record that other parts of the system depend on, with no validation step in between.
- Nobody can explain, in plain language, why the system made a particular decision about a user, a customer or an employee.
- A permission, approval or publication state can be changed by the model rather than by a rule or a person.
- The same input can produce different state on different days, and the product has no way to notice.
- When the model is unavailable, the workflow stops entirely, because the model was the workflow.
- The team cannot describe what the model's job is in one sentence, because its job is "everything".
How we draw the boundary
Here is the process we actually follow when we design an AI feature, whether it is for one of our products or a client's.
First, write down the workflow as states and transitions without any AI in it at all. Where does the work start, what are its stages, who moves it between stages, and what has to be true for each move to be allowed? If you cannot describe the workflow without AI, you do not understand the workflow yet.
Second, find the steps where a person is doing interpretation work that is slow, repetitive or inconsistent. Reading, categorising, drafting, extracting, summarising. Those are the candidates for AI assistance.
Third, for each candidate, define the job in one sentence, the input the model gets, the structured output it must return, and the validation the application performs before using it. If the output would change a consequential state, insert a human confirmation or a rule in between.
Fourth, decide what happens when the model fails. Not if. When. The answer should always be that the workflow stays intact and the user is told something useful.
Fifth, build the evaluation. Collect realistic inputs, define what a good output looks like for this specific job, and measure it. Measure it again every time the prompt, the model or the context changes.
That process is slower at the start than "just let the AI handle it". It is faster by the time you have real users, because you are not rebuilding the product around every failure you did not plan for.
If you cannot describe the workflow without AI in it, you do not understand the workflow yet.
This is not about distrusting AI
I want to be clear about something, because this argument is sometimes heard as scepticism about AI. It is the opposite.
We use AI heavily. Our products depend on it. The reason we keep it on the edge is that we want to use it as much as possible, and the only way to do that safely is to make sure a wrong output cannot do much damage. A model that is allowed to interpret freely inside a system that validates its work can be used everywhere. A model that is allowed to change state directly has to be used cautiously, because every call is a risk.
Constraining where the AI acts is what lets you be generous about how often it acts. The products that use AI most effectively are not the ones that gave it the most authority. They are the ones that gave it the clearest job.
What this means for your product
If you are planning to add AI to a workflow, or build a product around it, I would suggest three things.
Separate interpretation from consequence before you write any code. List the decisions in the workflow and sort them. The interpretation decisions are where AI will help you most. The consequence decisions are where it will hurt you most if it goes wrong.
Design the boundary as a feature. The input contract, the output structure, the validation, the fallback. This is the real engineering in an AI product, and it is what makes the AI usable.
Keep people where people belong. Near durable truth, near irreversible actions, near judgments about other people. Not because the model cannot form an opinion, but because someone has to be accountable for the decision, and a model cannot be.
Use AI where intelligence helps. Use software rules where certainty matters. Everything we build follows from that one sentence.
Related work
Related services