Why My AI Chatbot Gives Different Answers: My Troubleshooting Checklist

Illustration of a laptop with blank chat bubbles, a magnifying glass and three checkmarks

When an AI chatbot gives a different answer to the same question, I start by checking whether the answer is actually wrong. Different wording can be harmless. A missing requirement, an invented fact, or a recommendation that contradicts the supplied information deserves investigation. None of those observations, by itself, proves that the service changed its underlying large language model, or LLM.

My approach is to save one failing example, repeat a small task in a fresh conversation, and change one condition at a time. Then I separate what I can observe about the app from what I can only infer about its model. That gives me a useful next step: repair the task context, check a setting, avoid an unreliable answer, or send support a reproducible report.

The checklist below is designed for everyday PC work such as comparing files, following troubleshooting instructions, and summarizing notes. The sample task is a constructed exercise with an answer key, not a benchmark result or a claim that I tested particular commercial models.

First, I define what “different” means

I would not investigate a model change because one response says “open Settings” and another says “go to Settings.” Both might accomplish the same job. I would investigate if one response ignores the instruction to preserve files, invents an application setting, or recommends an action that the question explicitly ruled out.

Before repeating a question, I write down the condition that makes the answer usable. For a summary, that might mean retaining three dates and adding no unsupported facts. For a file comparison, it might mean selecting exactly the entries that satisfy the stated rules. A polished explanation does not compensate for getting that condition wrong.

This follows the useful principle in Anthropic’s guide to defining success criteria and evaluations: decide on specific, measurable criteria that fit the task. I apply that principle on a small scale here. I am diagnosing my own workflow, not trying to rank every model.

What changed?What I check firstWhat the observation establishes
Wording or paragraph orderWhether the same facts and requirements survivePossibly just another valid response
Facts or selected itemsThe supplied evidence and a written answer keyA task error if the answer contradicts them
Remembered detailsWhich details are present in the current conversationA context problem worth isolating
Speed, tone, or response lengthVisible mode, tools, instructions, and service conditionA product behavior change; its cause remains open
A displayed model labelWhere the label came from and when it appearedWhat that interface reports, within its own limits

I save the failing example before changing anything

A useful comparison needs something more precise than “it was better yesterday.” I keep the exact question, the relevant input, the full answer, and the mistake I can point to. If I still have yesterday’s successful answer, I preserve that too. Without it, I describe the earlier behavior as a recollection rather than a measured baseline.

I also record the app or website, time and time zone, visible model or mode label, and whether tools or attachments were involved. I note any custom instructions I know are active. Two windows can look similar while carrying different conversation histories, so a screenshot of the model menu alone is not enough for a fair comparison.

For shared evidence, I substitute a small fictional example for private material wherever possible. There is no need to attach a customer spreadsheet, a full personal conversation, or account credentials just to demonstrate that a chatbot selected the wrong filenames. I verify that the simplified example still shows the issue before sending it.

I use a tiny task whose answer I can check myself

Here is a PC-oriented exercise I would use when an assistant starts overlooking conditions. It tests whether the reply follows a few explicit rules. It does not measure general intelligence, deep reasoning, or model identity.

Use only this fictional file list. Return the names of files whose extension is .pdf, whose size is at most 10 MB, and whose modified date is on or after 2026-09-01. Every condition must be true. Return only the matching filenames, one per line. Do not add explanations.

notes.pdf | 4 MB | 2026-09-02
budget.xlsx | 2 MB | 2026-09-03
archive.pdf | 14 MB | 2026-09-04
manual.pdf | 8 MB | 2026-08-30
guide.pdf | 10 MB | 2026-09-01

The answer key contains notes.pdf and guide.pdf. The latter tests both inclusive boundaries: “at most” includes 10 MB, and “on or after” includes September 1. I did not request an output order, so either ordering is acceptable.

I score correctness and presentation separately. Selecting both required files and no others passes the selection check. Returning only filenames passes the format check. An explanation added to otherwise correct filenames is a formatting failure; selecting archive.pdf is a substantive error. Keeping those categories separate prevents me from calling every cosmetic change a loss of capability.

This exercise only becomes relevant if it resembles the failure I am investigating. If the problem concerns summaries, I would instead use a short fictional paragraph with a few facts that must survive. If the problem concerns code, I need a small example with a checkable expected result. Passing the file exercise cannot establish that an assistant will handle an unrelated task correctly.

I compare the old conversation with a fresh one

I first try the unchanged task in the conversation where the problem appeared. Then I open a new conversation in the same app, keep the visible model and mode the same where possible, and paste the identical task. I do not rewrite the question while also resetting the conversation, because that would change two conditions at once.

A fresh conversation is a practical control, not a guarantee of a completely blank configuration. Depending on the product, personalization, project instructions, or connected tools may still apply. I record the settings I can see and leave unknown settings marked unknown. I do not pretend to control hidden application behavior.

There is a research reason to take context seriously. The 2023 paper Lost in the Middle: How Language Models Use Long Contexts found that the position of relevant information affected performance on the tested retrieval and question-answering tasks. That historical result supports checking context; it does not prove that every current model gets worse as a chat grows.

If the fresh conversation succeeds and the old one fails, I treat that as a reason to examine accumulated instructions and irrelevant material. I can make a short handoff containing the current goal, confirmed facts, constraints, and unresolved question. I review that handoff myself before using it, because an automatically generated summary could preserve the very mistake I want to remove.

If both conversations fail, I check the task wording and answer key next. A shared failure can expose an ambiguous instruction or a limitation on this particular task. It still does not tell me which model generated the reply.

I change one visible setting and keep every result

Next I inspect the settings that actually matter to the failure: the selected mode, available tools, attachments, and custom instructions. I use the controls the application exposes. If it offers no setting for a particular behavior, I cannot claim to have held that behavior constant.

For the fictional file task, I do not need browsing or external documents. For a question about current software documentation, access to fresh sources could matter a great deal. Comparing a tool-enabled answer with an answer that had no source access would test two different workflows, even if their model labels matched.

I repeat each condition a few times when the result is inconsistent and save the failures as well as the successes. Three attempts per condition can be a manageable initial check, but it is not a statistically conclusive sample. I stop early if the problem is already reproducible and I have enough evidence for the next action.

ConditionKept unchangedRecord for each attempt
Original conversationExact task and visible modeSelected files, format, specific mistake
Fresh conversationExact task, app, visible modeThe same checks, including failures
One setting changedFresh conversation and exact taskWhich setting changed and the resulting checks

I avoid using “try again” as my only test. A retry inside the same conversation adds another instruction and may include feedback about the previous answer. That can be useful for getting work done, but it is a different comparison from submitting the original task under the same starting conditions.

I keep model-identity clues separate from task accuracy

If the practical checks still leave a model question, I look for information attributable to the service: a visible model label, documented version details, or a provider’s response metadata if I already have access to it. I do not ask a chatbot to certify its own identity and treat its generated sentence as an independent record.

A label tells me what that interface reports. To establish what ran behind an intermediary, I would need evidence with appropriate provenance from the service. A response becoming shorter, slower, or more formal does not provide that provenance. I would report the observed change without accusing a provider of secretly substituting a model.

For an additional clue when metadata is unavailable, I would use the llm checker at WhatsMyLLM. It compares replies to its supplied numerical challenge instructions with a reference bank. I would follow those instructions exactly, use a fresh chat for each, and paste the unedited replies rather than an ordinary essay or a self-description.

I would interpret its output as resemblance to a listed model, not verified identity. Its stated limitations include the possibility that an unlisted version resembles a listed relative. It does not establish the provider or subscription behind a service, and it does not measure whether the chatbot can complete my file task. A fingerprint result can add context to a report; it cannot replace the task checks above.

I use the evidence to choose a next step

If a clean conversation restores correct answers, my immediate action is to rebuild the working context carefully. If changing one visible setting restores the expected behavior, I record that setting and verify it on another relevant example. Either outcome can improve the workflow without resolving every question about the model.

If the assistant repeatedly fails a clear task under recorded conditions, I stop relying on it for that task until I have a dependable check. For consequential PC actions, that means verifying commands and file operations before executing them. I would not keep regenerating answers until one looks reassuring and then discard the contradictory results.

If I need help from the service, I use a short report with enough detail for someone else to reproduce the problem:

  • Environment: app or site, date, time zone, visible model and mode.
  • Input: the smallest non-sensitive question and data that still fail.
  • Expected result: the exact requirement or answer key.
  • Observed result: the full reply and the specific discrepancy.
  • Comparison: old versus fresh conversation, settings changed, and all attempt outcomes.
  • Remaining uncertainty: anything I could not inspect or hold constant.

My stopping rule is practical: once I can explain the failure well enough to fix the workflow, avoid an unreliable action, or file a reproducible report, I have something useful. I do not need to prove a hidden model switch to establish that an answer is wrong—or to make my next interaction more reliable.

Leave a Reply

Your email address will not be published. Required fields are marked *