Skip to content
TokIQ
Techniques

Chain-of-Thought Prompting Explained: Does It Still Work?

Chain-of-thought prompting asks a model to reason step by step before answering. See when it helps, when reasoning models make it redundant, and how to use it.

TokIQ Editorial6 min read
In this article
  1. Why does writing out the reasoning help?
  2. A concrete before and after
  3. Does "think step by step" still work with reasoning models?
  4. Should the model reason before or after the answer?
  5. How do you separate reasoning from the final output?
  6. When is chain-of-thought a bad idea?
  7. Related techniques worth knowing
  8. The short version

Chain-of-thought (CoT) prompting is asking a language model to work through intermediate steps in writing before it commits to an answer. It helps most on problems with several dependent steps, like arithmetic, logic and multi-constraint planning, and helps little on simple tasks. Newer "reasoning" models do this internally, which makes explicit step-by-step instructions less necessary than they were a few years ago.

The technique has a clear origin. In 2022, Wei et al. published "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models", showing that few-shot examples which included worked reasoning, not just answers, sharply improved large models on math word problems and similar benchmarks. Shortly after, Kojima et al. ("Large Language Models are Zero-Shot Reasoners", 2022) found that simply appending "Let's think step by step" produced a similar effect without any examples. That phrase has been copied into a million prompts since.

Why does writing out the reasoning help?

A model generates text one token at a time, and each token is conditioned on what came before. If it has to output the final answer immediately, all the work has to happen in a single step. If it first writes "the train leaves at 14:10, the trip takes 2 hours 35 minutes, so it arrives at 16:45", each intermediate result becomes part of the context for the next step. The written steps act as scratch space.

That explanation also tells you where CoT will not help. If the task does not decompose into steps, there is nothing for the scratch space to hold. Asking a model to "think step by step" about which of two synonyms sounds friendlier mostly produces a longer answer.

A concrete before and after

Here is a problem that trips up models answering in one shot:

A meeting room is booked 09:00–10:30 and 13:00–14:00.
A cleaning crew needs 45 minutes, must finish before 12:00,
and cannot start before 09:15. List every valid start time
on the quarter hour.

Answering immediately, a model may list 09:15, which overlaps the booking. A version that asks for the work first:

A meeting room is booked 09:00–10:30 and 13:00–14:00.
A cleaning crew needs 45 minutes, must finish before 12:00,
and cannot start before 09:15.

First, write out the free windows between 09:15 and 12:00.
Then check which quarter-hour start times fit a 45-minute
slot inside those windows. Finally list the valid start times.

Notice that this is better than a bare "think step by step". It tells the model which steps matter. Generic CoT lets the model choose its own decomposition; guided CoT gives it the right one. When you know how the problem should be broken down, say so.

(The answer, for the record: free window 10:30 to 12:00, so 10:30, 10:45 and 11:00, with 11:15 excluded because the slot must finish before 12:00. Writing the steps out makes that last boundary hard to miss.)

Does "think step by step" still work with reasoning models?

This is where advice from 2023 goes stale. Several model families now include reasoning modes or dedicated reasoning models that generate a hidden or summarized chain of thought before the visible answer. OpenAI's o-series, Anthropic's extended thinking, and Gemini's thinking models are all examples of this approach.

With those models, a few things change:

  • Adding "think step by step" usually does little, because the model already spends tokens reasoning.
  • The vendors' own guidance tends to recommend high-level instructions over prescribing exact steps. Telling a reasoning model precisely how to think can constrain it and sometimes makes results worse.
  • You often control reasoning through an API setting (a reasoning effort level or a thinking token budget) rather than through prompt text.

So the practical rule is: on a standard chat model, explicit CoT is still a useful tool for multi-step problems. On a reasoning model, describe the goal, the constraints and what a correct answer must satisfy, and let it plan. Check the provider's documentation for the specific model you use, because this area changes quickly.

Should the model reason before or after the answer?

Before. Always before.

This is one of the most common mistakes in structured prompts:

{
  "answer": "approve",
  "reasoning": "The applicant meets the income threshold and..."
}

If the answer field comes first, the model commits to "approve" and then writes a justification for a decision it already made. The reasoning becomes a rationalization. Reverse the order so the reasoning is generated first and the answer can depend on it:

{
  "reasoning": "Income is 52,000, threshold is 50,000, so income passes. Debt ratio is 0.46, limit is 0.40, so it fails...",
  "answer": "reject"
}

The same applies to plain text: "Explain your reasoning, then give the final answer on the last line" beats "Give the answer and explain why."

One caveat worth stating honestly. The written chain of thought is not guaranteed to be a faithful account of how the model reached its answer. Research on CoT faithfulness has found cases where the stated reasoning omits factors that actually influenced the output. Treat the reasoning as a useful work product, not as an audit log.

How do you separate reasoning from the final output?

In an application, users usually should not see the scratch work. They want the answer. Two reliable patterns:

Tags in plain text.

Work through the problem inside <thinking> tags.
Then give only the final answer inside <answer> tags.
The <answer> must be a single number with no units.

Your code extracts whatever sits inside <answer> and discards the rest. Anthropic's documentation recommends this tag-based pattern for Claude, and it works well across models.

Separate fields in structured output. If you already request JSON, add a reasoning field before the answer field, as shown above, and only display answer. Our guide on reliable JSON from LLMs covers how to enforce that structure.

With reasoning models that think internally, you typically get the final answer directly, sometimes with an optional reasoning summary from the API. You do not need tags for the hidden part.

When is chain-of-thought a bad idea?

It is easy to overuse. Skip it when:

  • The task is a lookup or a simple transformation. Translating a sentence, fixing grammar, extracting an email address. CoT adds cost and occasionally talks the model out of the obvious right answer.
  • Latency matters more than marginal accuracy. Reasoning tokens are generated tokens. A chatbot that pauses for six seconds to deliberate over "what are your opening hours" is a worse product.
  • You are classifying with crisp rules. Clear definitions and a couple of examples usually beat reasoning here. See zero-shot vs few-shot prompting.
  • You would show the reasoning to end users without review. Intermediate steps can contain half-formed or wrong claims that the final answer later corrects.

Self-consistency (Wang et al., 2022) samples several chains of thought and takes the majority answer. It improves accuracy on problems with a single correct answer, at the cost of running the model multiple times.

Decomposition across calls. Instead of one long prompt that does everything, split the problem: one call extracts the facts, a second reasons over them, a third formats the result. Each prompt is easier to test, and you can see exactly where a failure happens. This is often the more robust version of CoT in production systems, and it leads naturally into tool calling and agents.

Asking for verification. After the model produces an answer, a second prompt asks it to check that answer against the original constraints. It catches some errors, though a model checking its own work shares its own blind spots.

The short version

Chain-of-thought gives a model room to work through dependent steps, and that room matters most for math, logic and planning. Tell it which steps matter when you know them. Put reasoning before the answer, never after. Hide the scratch work from users with tags or separate fields. And on reasoning models, describe the goal rather than scripting the thinking.

If you want to practice spotting when reasoning helps and where it is placed wrong, the reasoning topic walks through those decisions one at a time.

Frequently asked questions

What is chain-of-thought prompting?

Chain-of-thought prompting asks a language model to write out intermediate reasoning steps before giving its final answer. It was popularized by Wei et al. (2022), who showed large gains on math and logic benchmarks.

Does "let's think step by step" still work?

On standard chat models it can still help with multi-step problems. On reasoning models that already think before answering, adding it usually changes little, and vendors generally advise against micromanaging their internal reasoning.

When should I not use chain-of-thought?

Skip it for simple lookups, classification with clear rules, or short rewrites. It adds tokens and latency, and on easy tasks the extra text can introduce errors instead of preventing them.

How do I hide the reasoning from the final output?

Ask the model to put its reasoning inside one tag, such as <thinking>, and the final answer inside another, such as <answer>, then show only the answer to users. With structured output, use a separate field for the reasoning.

  • #chain of thought
  • #reasoning
  • #think step by step
  • #reasoning models

Now practice it

TokIQ turns prompt engineering into short quizzes, with an explanation for every answer.

Coming soon onApp StoreComing soon onGoogle Play

Write better prompts, a few questions a day.

Short quizzes on real prompting decisions, with an explanation for every answer. Free to start on iPhone and Android.

Coming soon onApp StoreComing soon onGoogle Play