← All writing

The Benchslap, and How I Work With AI

A lawyer got sanctioned for filing a brief with AI-generated citations to cases that don't exist. The failure wasn't the model, it was a workflow with no review step. Here is the paper-to-outline-to-session process that keeps a professional behind every output.

Someone I follow on LinkedIn coined the term “benchslap” a few months ago, and I loved it immediately. The author was referring to a court imposing a fine on an attorney for submitting a brief using generated language that included citations to caselaw that didn’t exist and several other problems.

The word is funny. The situation is not. That’s a malpractice problem wearing a technology costume.

The Workflow Failure

Every profession I know has a version of this requirement: work product doesn’t leave without a responsible professional behind it. First-year associates don’t send their initial case analyses to clients. Nobody debates whether that step should exist. It exists because someone’s name and license are attached to what goes out the door.

AI output has to be verified as well. Gen AI tools are often fast and the output can be good, but it is occasionally wrong in ways that are genuinely hard to catch. The benchslapped lawyer didn’t have a bad model. They had a workflow with no review step.

That leaves the harder question: what does a workflow with the right structure look like? Here’s mine.

One Question, Three Answers

The question I hear from practitioners trying to figure out how to work with AI is usually some version of “should I use it or not.” That’s the wrong question. It treats the relationship like a one-way door you either walk through or decline.

The question you’re actually addressing is: where am I right now in this work, and what do I need?

I’ve found three good approaches that can be stacked and combined as needed. In one, you’re staying in your own thinking, still forming your position, and AI has no role yet. In another, you’re using the AI as a mirror, describing your direction and letting it run to stress test your reasoning before you commit to it. In the third, you’re in structured collaboration, working with a thought partner who interrogates something you’ve already built.

The mirror mode is one I get tremendous value out of, and I think it’s one of the most important techniques in this kind of work. After I’ve developed an idea, I ask the model to use the existing high-level outline and context and turn it into a complete outline. The purpose is to find out whether your thinking holds up before you invest in it. If you describe your goal and your reasoning clearly, the AI produces something coherent you can refine. If your reasoning has a gap, or your direction is murkier than you thought, the output reveals it. You find the problem cheaply, at the outline stage, before you build anything on top of a flawed premise.

The clearest signal is what happens when the output wanders: you find yourself reaching for an analogy to redirect it. That analogy is the thing you should have said in the first place. Your thinking needed it. The AI just helped you find it.

The AI is a mirror. Blurry output means blurry input. When it’s not giving you something useful, look at what you gave it, not at the model.

The workflow that led up to the benchslap wasn’t just a missed review step. It was missing the creation of a structure that maximizes the value of AI assistance while mitigating the risk of AI malfunction. The lawyer didn’t appear to have structured or stress tested anything. They went straight to “done, hand it off,” and the model had incomplete information to work from.

Starting on Paper

My process starts on paper. Not metaphorically. Literally. Pen, notebook, no structure imposed. The goal isn’t to produce anything. I’m just determining whether I even have a point worth pursuing.

I write until I recognize my thesis. Then I think about how I want to deliver my message. A paragraph, half a page, sometimes a diagram. The artifact doesn’t matter. The point does.

I do this before I open a session. The AI has no role here and shouldn’t. You can’t get useful output from a model without doing this work first. You can’t evaluate a response to a position you don’t hold.

Letting the Model Run

When I have a direction, a clear enough goal and some reasoning but not necessarily a clean outline, I open a session, describe what I’m trying to do, and let the AI take a run at the outline structure. I’m not asking for a draft. I’m asking it to build out the logical scaffolding from what I gave it so I can see whether that scaffolding holds.

The purpose is to stress test my thinking before I commit to a direction. If the AI produces a coherent structure from what I gave it, my reasoning held up. If it wanders or misreads the goal, my reasoning had a gap and the output shows me where.

This is how I avoid building the wrong thing. The mistake surfaces at the outline stage, where it costs nothing, instead of after I’ve invested time drafting something built on a flawed premise.

Sometimes I use an analogy to redirect the output when it goes sideways. My description was missing something, and the AI’s misfire helped me find it. That analogy usually turns out to be the thing I should have said in the first place.

This mode can also produce position changes. When the AI’s interpretation of a vague prompt reveals something unexpected about what you actually think, pay attention. That’s information about your own reasoning, not the model’s.

Shifting Into Structured Collaboration

When the mirror is giving back something recognizable, and the output feels like a response to what I actually meant, I transition into the next phase.

I open a session with a precise first sentence defining the goal. Then operating instructions for the model. Then I tell it explicitly that we’re iterating the outline before any drafting begins. The AI’s job at this stage isn’t to generate content. It’s to interrogate the structure, find the gaps I can’t see from inside my own thinking, and push back when I’m being vague.

Every decision is mine. When the model identifies a gap, I decide whether it matters and how to address it. A polished draft built on a weak structure is worse than a rough draft of a solid one. The outline iteration is the highest-leverage moment in the process, and it’s low-stakes because nothing has been written yet. I generally do about five iterations on an outline, and about two passes on those iterations performing manual line-by-line revisions myself. Fewer iterations means I brought better content and analysis. More iterations, and more manual revision passes in particular, indicate I was less clear going in.

When the structure holds, I move to draft. Same dynamic: I bring the draft, the model pushes it toward what it was trying to be, I judge what to keep. At some point the mode shifts again, less structure-and-direction work, more execution and precision. The decisions stay mine throughout. What changes is what I’m asking for.

Building Infrastructure

The paper-to-outline-to-session workflow works for a single piece. The compounding version is building reference infrastructure: voice documents, brand context files, topic reference material.

Instead of describing my position from scratch every session, I build documents that encode my standards once. A voice observations document. A brand context file with my audience, my messaging pillars, what I don’t say. Now the AI isn’t guessing what edits I’m going to make. It generates content I need to change less, because the model has a document that says what sounds like me, written and reviewed by me.

This is human-in-the-loop thinking applied at the system level. The judgment was put in deliberately, once. It governs every session after that. This piece was written with a brand context document loaded at the start of the session that provides context to the model about my writing style, editing expectations, workflow cadence, whatever I think will help the AI understand how I write and edit. The AI knew my voice, my audience, and what I don’t say before I typed the first sentence about the benchslap.

The readers who build this infrastructure get a qualitatively different result than the ones who don’t. It compounds.

How I Explain the Relationship

I now say “Claude was my editor on this piece” rather than “AI-assisted” or “written with AI.”

An editor makes you better. A ghostwriter replaces you. Most people don’t yet have a clean model for the difference, and the framing matters, both for honesty and for how you think about the work yourself. The writing is mine. The structural decisions, the observations, the voice are mine. The model’s job was the same as any good editor’s: don’t let me get away with imprecision.

The output chain has a professional behind it at every step. That’s what was missing in the benchslap.

If you’re trying to integrate AI into professional work, the first question isn’t “what can it do.”

It’s: where am I right now in this work, and what do I need?

Answer that before you open a session, and the benchslap problem largely takes care of itself.

Claude was my editor on this piece.

See the platform in action.

Free. Dan applies the Rote methodology to your situation and delivers structured findings within one week.