This one is a sliding scale, rather than a straightforward answer.

I also think it’s the wrong question.

I think the right question is: how much of the response from the LLM do I need to check?

An even better question is: what am I letting the AI take over, and what am I leaving under my own control?

Let’s start with how much of the response from your LLM you should be checking.

Risk Assessment

When using AI for market research, one of the primary questions you can ask yourself to determine how much of the output you should check is, "What level of risk do I have if the output is wrong?"

Low Risk

Low-risk tasks are those where the output is either going to be reviewed anyway or being wrong is going to cost you all of 10 minutes. Draft surveys or discussion guides, ideas for talk titles, brainstorming.

In these cases, the check is either built-in, or you probably don’t need one.

Medium Risk

Medium risk tasks are those where being wrong might cost you an extra review cycle, but isn’t going to cost you a contract or customer relationship or your reputation. Coding open-ends, theme analysis, reviewing a report you already wrote and just want to check against the data to see if you missed something.

In these cases, do a spot-check.

High Risk

You’ve probably guessed this by now, but this is where being wrong would cost you a customer relationship or your reputation. Writing entire reports, creating presentations to be delivered to an audience, writing articles for publications. Example: you upload a csv or Excel file and tell the AI to analyze the data to answer a specific business question, then create a set of slides with charts and data

In these cases, check everything.

If I’m Checking Everything, Then What’s the Point?

This is a question I have wrestled with myself often, and I recently landed on this answer.

In the pre-LLM scenario, you were analyzing data, reading transcripts, manually coding open-ends, creating a thematic analysis of the transcripts, and then taking notes, organizing your findings, and drafting a report, creating the charts, writing headlines, etc. Then you’d likely send what was created to someone else to check - the QC run.

In the LLM scenario, you’re reviewing the data and reading the transcripts, but you now have the option of having the LLM code the open-ends, do a thematic analysis of the transcripts, draft the report, create the charts, and write the headlines - and you are the one doing the QC work.

This is where the time-savings happens. Even if you’re checking every data point and editing the report by moving things around and writing stronger recommendations, you bypassed the initial drafting stage.

Pay Attention to What Thinking You Hand the LLM

I initially said the better question is: what am I letting the AI take over, and what am I leaving under my own control?

I think this is the key question that we need to be asking more often.

It’s easy to stop checking the output from an LLM when 90% of the time, the QC uncovers nothing. At that point, it becomes easy to essentially hand over the data analysis, theme analysis, and reporting entirely over to the LLM and just do a spot check here and there and call it good.

That’s where we’re handing our control and our judgement to the LLM.

Think about the scenario if you replace “LLM” with “peer.” It’s easy to stop checking the report written by a peer when 90% of the time, the QC phase uncovers nothing. But does that mean you stop the QC process? Of course not! Humans fail randomly - a bad day, a tired day, a distracted-by-life-events day. Where there used to be no errors, suddenly, there are errors. The report is still going to a client, and getting something wrong is still a bad look on you. So you’re still going to QC.

The same should apply for LLM output that’s going to a client. LLMs fail unpredictably, too, but because it’s a machine, we’re lulled into a sense that they are predictable. The problem is that models change, and each change introduces chances for different errors. And even static models have error rates - the risk doesn’t go away because nothing else changed. That error rate means we should be checking the work, especially when it matters most.

I’m launching pilot beginner and intermediate cohorts at half price in exchange for feedback to improve the courses for future students. Classes are held weekly, recorded for those who can’t make it, and all include hands-on application of principles learned using market research work examples.

Just hit reply to this newsletter if you’d like to register for the September cohorts. There are limited spots available for each cohort, so reply soon!

Upskill your AI skills.