A firm buys seats for an assistant. Somebody asks it a real question about a real client. The answer comes back confident, generic, and wrong in a way that takes a minute to notice.

The usual conclusion is that the model is not good enough yet. That is almost never what happened.

The boundary nobody draws on the whiteboard

A source it cannot reach produces no error. The answer arrives without it.
A source it cannot reach produces no error. The answer arrives without it.

An assistant can only use what it can reach. When it is sold to a firm, what it can reach is usually one thing: the productivity suite it ships with. Mail, calendar, the documents in that suite's own storage, the chat history.

Meanwhile the answer to the question lives somewhere else. The engagement record is in the practice management system. The prior year workpapers are in the document management system. The current figures are in the accounting platform. The agreement about how this client's inventory gets treated was made in a meeting and written into a memo that lives in a folder whose name made sense to one person in 2023.

The assistant can see the email where somebody mentioned the memo. It cannot see the memo.

So it does what a capable person with partial information does. It produces something that sounds like the shape of an answer, drawn from general knowledge rather than from your firm. The confidence is a property of how these systems write. The emptiness is a property of what they were allowed to read.

Why this reads as a model problem

Three things make the boundary hard to see from the outside.

The failure is silent. A system that cannot reach a source does not announce it. There is no error, no gap in the page, no note saying "I could not open the document management system". The answer simply arrives, formed and plausible.

The demonstrations are honest and misleading at once. A vendor demonstration works because everything it touches is inside the one system it can see. Summarise this thread. Draft a reply. Find the attachment from last Tuesday. All of those are genuinely useful and all of them stay inside the boundary. The question that crosses it is the one a partner asks on day three.

The good answers and the bad ones look identical. A summary of an email thread and an invented client history are rendered in the same tone, the same length, the same formatting. Nothing in the presentation tells you which one you are reading.

That last property is the one worth sitting with. A system that failed loudly would be a nuisance. A system that fails plausibly costs you the time of whoever checks it, and it costs you more if nobody does.

What changes when the boundary moves

The fix is unglamorous. The assistant needs to be able to reach the systems that hold the answer, and it needs to say where each part of its answer came from.

That second half does more work than it looks. When every claim carries a source, three things become true at once. A reviewer can open the original and check it in seconds. A claim with no source is visible as a claim with no source. And the system can be honest about coming up empty, because "I found nothing in the document system about this" is a sentence it can now say truthfully.

A firm that has this stops asking whether the answer is right and starts asking whether the citation supports it. That is a much faster question to answer, and it is one that a junior person can do reliably.

What the seats are still good for

None of this is an argument for cancelling the subscription. The assistant your firm already pays for is genuinely useful inside its boundary, and that boundary contains a lot of the working day.

Drafting and rewriting. Summarising a long thread before a call. Turning rough notes into something sendable. Reformatting. Finding the message you half remember. Explaining an unfamiliar concept at whatever level of detail you ask for. All of that works, none of it needs access to your practice systems, and a firm that uses the seats well for those things gets real time back.

The line is specific. Ask it to do work that depends on what your firm knows about a particular client, and it is answering from the part of the world it can see, which does not include your firm.

The question to ask a vendor

When somebody proposes an assistant for your firm, one question sorts most of it out.

Which systems will it be able to read, and how will an answer show me where it came from?

The first half tells you where the boundary will sit. Listen for specific systems named, and for how each connection is made. The second half tells you whether you will be able to check the work, and a vendor who has thought about citation has usually thought about accuracy.

If the answer to either half is vague, the boundary is going to end up wherever the software happens to put it. That is worth discovering in the meeting rather than on day three, when a partner asks the first question that crosses the line and gets back something fluent and hollow.