Skip to content
TokIQ
Techniques

How to Write a System Prompt (With a Full Working Example)

A system prompt sets an AI assistant's role, rules and format for a whole conversation. Learn what belongs in it, how to order it, and see a full example.

TokIQ Editorial5 min read
In this article
  1. What belongs in the system prompt vs the user message?
  2. Do roles and personas actually help?
  3. How should you order a system prompt?
  4. A full system prompt example
  5. Why explain the reason behind a rule?
  6. How do you test a system prompt?
  7. What should never go in a system prompt?
  8. Does the system prompt work the same on every model?

A system prompt is the standing set of instructions that tells a model who it is acting as, what it should and should not do, and how to format its answers, for every turn of a conversation. Write it like a briefing for a capable new colleague: the context of the product, the audience, the rules that matter, and how to handle the cases where things go wrong.

Most chat APIs (OpenAI, Anthropic, Google) accept a system or developer message separate from user messages, and models are trained to weight it more heavily. That makes it the natural place for anything that should stay true across the whole conversation.

Key takeaways

  • Put stable rules, context and format in the system prompt. Put the specific task and data in the user message.
  • A role helps, but context about the product and audience does far more than a job title.
  • Explain why a rule exists; models apply reasoned rules more sensibly than bare commands.
  • Write explicit behavior for edge cases: missing information, off-topic requests, angry users.
  • Never put secrets in a system prompt and never rely on it alone for security.

What belongs in the system prompt vs the user message?

The simplest test is: will this be true for every request? If yes, system prompt. If it changes per request, user message.

System promptUser message
Product context and audienceThe user's actual question
Role and toneThe document to summarize or analyze
Rules and boundariesRetrieved search results for this turn
Output formatPer-request options ("make this one shorter")
How to handle edge casesCurrent date or user profile, if it varies

Retrieved documents are an interesting case. Some teams put them in the system prompt; we prefer the user turn, because retrieved text is untrusted data, and the system prompt should be the place for instructions you wrote yourself. That separation also makes prompt injection a bit easier to reason about.

Do roles and personas actually help?

Somewhat. "You are an experienced pediatric nurse" shifts vocabulary, depth and caution in a useful direction. Anthropic's documentation explicitly recommends giving Claude a role via the system prompt, and the effect is real.

But the role is the weakest part of a good system prompt. Compare:

You are a world-class customer support expert.
You answer support questions for Ledgerly, a bookkeeping app used by
freelancers and small agencies in the US. Most users are not accountants.
They usually write in when an invoice export fails or a bank sync stops.

The second one never claims expertise. It gives the model what an expert would actually know: who the users are, what they struggle with, what the product does. The model will sound more competent with that than with any adjective.

Personas with elaborate backstories ("You are Max, a 34-year-old former barista who loves helping people") rarely improve accuracy. They can make sense for entertainment products. For utility products, they mostly add tokens.

How should you order a system prompt?

There is no single correct order, but this one works well and is easy for humans to maintain:

  1. Who and what. One short paragraph: the product, the assistant's job, the audience.
  2. Context. Facts the model needs: features, policies, what the product cannot do.
  3. Rules. Behavior constraints, each with a reason when the reason is not obvious.
  4. Edge cases. What to do when information is missing, the request is off-topic, or the user is upset.
  5. Output format. Length, structure, markdown or not, language.
  6. Examples (optional). One or two short exchanges showing tone and format.

Put long reference material (a product manual, a policy document) in clearly delimited blocks, for example XML-style tags. Anthropic's guidance recommends tags like <policy> for this, and they help with any model because they make boundaries explicit.

A full system prompt example

Here is a complete system prompt for a fictional bookkeeping app's support assistant. It is long on purpose; real ones are.

You are the in-app support assistant for Ledgerly, a bookkeeping app for
freelancers and small agencies in the US. Users are mostly designers,
developers and consultants. Most are not accountants and get anxious
about taxes, so be calm and concrete.

<product_facts>
- Plans: Solo (1 user) and Team (up to 10 users).
- Bank sync works with US banks through a third-party provider.
  Syncs run every 6 hours; users can trigger a manual sync in Settings > Banks.
- Invoices can be exported as PDF or CSV.
- Ledgerly does not file taxes and does not give tax advice.
</product_facts>

Rules:
- Only answer questions about using Ledgerly. If asked about something
  unrelated, say briefly that you can only help with Ledgerly.
- Do not give tax or legal advice, even if asked directly. Users may act on
  it and you cannot see their full situation. Suggest they ask a tax
  professional, and offer help with the Ledgerly side of the question.
- Never guess at account-specific data such as balances or invoice status.
  You cannot see the user's account. Tell them where in the app to look.
- If a feature is not listed in <product_facts>, do not claim it exists.
  Say you are not sure and suggest contacting support@ledgerly.example.

When a bank sync problem is reported:
1. Ask which bank, and when the last successful sync happened, if not stated.
2. Suggest a manual sync in Settings > Banks.
3. If that fails, suggest reconnecting the bank.
4. If it still fails, offer to hand off to a human agent.

If the user is frustrated, acknowledge it in one short sentence and move
to the fix. Do not apologize repeatedly.

Format:
- Plain text, no markdown headings. Numbered steps when giving instructions.
- Keep answers under 120 words unless the user asks for more detail.
- Reply in the language the user writes in.

A few things worth pointing out:

The tax-advice rule explains its reason ("Users may act on it and you cannot see their full situation"). With a reason, the model can generalize: it will also be careful about questions like "can I deduct my laptop?", which the rule never mentions by name.

The "do not claim features exist" rule is there because product assistants love to invent plausible settings menus. If you ship a support bot, this will be your most common bug, and grounding it in a facts block plus an explicit fallback is the fix. For larger knowledge bases, that facts block becomes retrieval; see what RAG is.

The escalation steps are numbered because order matters. Unordered rules about the same situation tend to get applied unpredictably.

Why explain the reason behind a rule?

Bare commands in capital letters ("NEVER mention competitors!!!") were a habit from older, less attentive models. Current models follow instructions closely, and shouting tends to make them overapply a rule. A model told "NEVER discuss pricing" may refuse to say whether a feature is on the free plan, which users reasonably need to know.

Compare that with: "Do not quote prices, because they vary by region and change often; link to the pricing page instead." The model now knows the boundary (specific numbers) and the intent (avoid stale information), so it can still answer "is export available on Solo?"

How do you test a system prompt?

Write down twenty to fifty realistic user messages before you change anything, including the nasty ones:

  • An off-topic request ("write me a poem").
  • A request for the forbidden thing, phrased politely and then phrased cleverly.
  • A message with missing information.
  • A message in another language.
  • An angry message.
  • Something your product genuinely does not support.

Run them all after every change. System prompts accumulate rules over time, and rules interact. Adding "always be concise" can quietly break the escalation flow that needed three steps. You only find out if you rerun the whole set.

What should never go in a system prompt?

Secrets. API keys, internal URLs, other customers' data. Users regularly get models to reveal their system prompts, and no wording reliably prevents it. Assume it is public.

Security enforcement you cannot back up in code. "Only issue refunds under $50" in a prompt is a suggestion. The refund API should enforce the limit. The prompt injection guide and the safety topic explain why.

Contradictions. Long system prompts written by several people tend to contain them ("always be thorough" in one section, "keep it short" in another). The model will pick one, and not consistently. Read the whole thing top to bottom every so often.

Does the system prompt work the same on every model?

Mostly, with differences in how strictly different models weight it, how they handle very long prompts, and what formatting they respond to best. Test on the model you ship with, and read that vendor's prompting guide. We cover what transfers in prompting ChatGPT vs Claude vs Gemini.

For practice on the specific decisions (what goes where, which role statement adds information and which only adds words), the role and context topic is a good place to start.

Frequently asked questions

What is a system prompt?

A system prompt is a set of instructions supplied to the model separately from the user's messages, usually by the developer. It defines the assistant's role, rules, tone and output format for the whole conversation.

What is the difference between a system prompt and a user prompt?

The system prompt holds stable instructions that apply to every turn, written by whoever builds the product. The user prompt holds the specific request and the data for this turn, often written by the end user.

How long should a system prompt be?

As long as it needs to be to cover real behavior, and no longer. Many production system prompts run from a few hundred to a few thousand words, but every rule should exist because a test case needed it.

Can users see or override the system prompt?

Models are trained to give system instructions priority, but that is not a guarantee. Users can sometimes extract or circumvent the system prompt, so never put secrets in it and enforce critical rules in code.

  • #system prompt
  • #roles
  • #personas
  • #LLM applications

Now practice it

TokIQ turns prompt engineering into short quizzes, with an explanation for every answer.

Coming soon onApp StoreComing soon onGoogle Play

Write better prompts, a few questions a day.

Short quizzes on real prompting decisions, with an explanation for every answer. Free to start on iPhone and Android.

Coming soon onApp StoreComing soon onGoogle Play