What is a model context window, in practice

Tools and infrastructure3 min read

A context window is the maximum amount of text, measured in tokens, that a model can read and reason over in a single call. Everything the model knows during that call, prompts, history, retrieved documents, tool outputs, must fit within it.

Tokens, not words

Models do not process characters or words directly. They process tokens, which are chunks of text produced by a tokeniser. A rough rule of thumb is that one token equals about three to four characters in English, or roughly 0.75 words. A 128,000-token context window holds something in the region of 90,000 to 100,000 words, though exact figures depend on the model and the tokeniser it uses.

Code tokenises differently from prose. Variable names, punctuation-heavy syntax and non-Latin characters can consume tokens faster than plain English text.

What lives inside the window

Every token in a single model call competes for the same space:

  • System prompt
  • Conversation history
  • Retrieved documents or chunks (see what is RAG)
  • Tool definitions and tool call results
  • The user's current message
  • Space reserved for the model's reply

In an agentic loop, tool outputs return into the context on every iteration. A multi-step task can fill a large window faster than engineers expect.

Why size is not the whole story

Longer context windows do not solve everything, for a few reasons.

First, cost. Most providers charge per token, input and output. Sending 100,000 tokens on every call is expensive at scale.

Second, latency. Larger payloads take longer to process, which matters in interactive or real-time applications.

Third, retrieval quality. Research and engineering experience suggest that models can lose track of information buried in the middle of a very long context. Placing the most relevant material near the start or end of the prompt tends to produce better results.

Fourth, the context window is ephemeral. Nothing in it persists between calls unless the application explicitly manages it. Memory across sessions must be handled at the application layer, not assumed from the model.

Practical consequences for engineers

Engineers working at the application layer regularly need to make decisions about context management:

  • Which parts of a conversation to keep and which to summarise or discard
  • How many retrieved chunks to pass in, and how to rank them
  • Whether to split a large document or route queries across smaller, focused calls
  • How to handle tool-heavy agentic tasks without hitting token limits mid-run

These are not configuration choices. They require understanding what the model actually does with the information it receives. An engineer who treats the context window as simply a larger bucket tends to produce systems that are slow, expensive, or unreliable under realistic workloads.

Context engineering is the discipline of making deliberate, reasoned decisions about what to put into a context window, how to structure it, and what to leave out.

Relationship to other components

A vector database exists partly because context windows are finite. Rather than loading an entire knowledge base into a prompt, retrieval systems find the most relevant chunks and pass only those. The quality of that retrieval directly affects what the model can reason over.

In agentic coding setups, the context window also determines how much of a codebase a model can hold in view at once. Engineers working on large repositories often need to be deliberate about which files, functions or diffs they include.

What we test for

Managing a context window well is something we watch for directly in how we assess engineers. In the how we vet AI-native assessment, we score prompt and context engineering on exactly this: whether an engineer sets up the task context so the tool produces usable output, rather than passing everything available and re-prompting when it fails. We also look at spec-driven development, because an engineer who breaks a vague ticket into well-scoped pieces is making the same kind of decision: choosing what the model needs, and leaving out what it does not.

Short answers

Does a larger context window mean better results?

Not automatically. Larger windows increase cost and latency, and models can lose track of information buried deep in long prompts. Careful selection of what enters the context often matters more than raw window size.

What happens when you exceed the context window?

The API returns an error, or the provider silently truncates the input, depending on implementation. Either way, content is lost. Engineers must handle this case explicitly in production systems, not treat it as an edge case.

Is context window size the same across all models?

No. It varies significantly by model and version, from a few thousand tokens to over a million in some cases. The relevant limit for your system is the model you are actually calling, not the largest number you have seen quoted.

Let's talk

Get a shortlist within five working days

You share the roles and the stack in a short form or a thirty-minute call. Within five working days you get named senior engineers to review, each with both scorecards.

Reviewed onClutch4.9 out of 5 from 36 reviews
ISO 27001
Certified

Book thirty minutes with Dale

The calendar is provided by HubSpot, which sets its own cookies. Load it here, or book on HubSpot's page.

Open booking page