GitHub Copilot Model Comparison for Coding

Prompting, Models & CustomizationAcademy lesson 73Cluster 6 · Lesson 10 of 12Intermediate17 min readVersion-sensitive
Published
Updated
Last technically verified
GitHub Copilot Model Comparison for CodingPrompting, Models & Customization10Intermediate/github-copilot/customization/model-comparison/

The current Copilot model inventory in full: 32 models from 7 providers, with what GitHub publishes about each and nothing it does not.

The full inventory

Model data last verified2026-09-01Model availability changes frequently.
Every model GitHub currently lists as supported
ModelProviderStatusGitHub groups it under
GPT-5 miniOpenAIGAGeneral-purpose coding and writing; Deep reasoning and debugging; Working with visuals
GPT-5.3-CodexOpenAIGAGeneral-purpose coding and writing
GPT-5.4OpenAIGANo task grouping published
GPT-5.4 miniOpenAIGANo task grouping published
GPT-5.4 nanoOpenAIGANo task grouping published
GPT-5.5OpenAIGADeep reasoning and debugging
GPT-5.6 LunaOpenAIGAFast help with simple or repetitive tasks
GPT-5.6 SolOpenAIGADeep reasoning and debugging
GPT-5.6 TerraOpenAIGAGeneral-purpose coding and writing
Claude Fable 5AnthropicGANo task grouping published
Claude Fable 5.1AnthropicGANo task grouping published
Claude Haiku 4.5AnthropicGAFast help with simple or repetitive tasks
Claude Opus 4.5AnthropicGANo task grouping publishedRetires 2026-09-01
Claude Opus 4.6AnthropicGANo task grouping publishedRetires 2026-09-01
Claude Opus 4.7AnthropicGADeep reasoning and debugging
Claude Opus 4.8AnthropicGANo task grouping published
Claude Opus 4.8 (fast mode)AnthropicPublic previewNo task grouping published
Claude Opus 5AnthropicGANo task grouping published
Claude Sonnet 4.5AnthropicGANo task grouping publishedRetires 2026-09-01
Claude Sonnet 4.6AnthropicGADeep reasoning and debugging; Working with visualsRetires 2026-09-01
Claude Sonnet 5AnthropicGANo task grouping published
Gemini 3.1 ProGooglePublic previewDeep reasoning and debugging; Working with visualsRetires 2026-09-01
Gemini 3.5 FlashGoogleGANo task grouping published
Gemini 3.6 FlashGoogleGANo task grouping published
Gemini 3.7 FlashGoogleGANo task grouping published
MAI-Code-1-FlashMicrosoftGAGeneral-purpose coding and writing; Fast help with simple or repetitive tasksRetires 2026-09-10
MAI-Code-1.1-FlashMicrosoftGANo task grouping published
Raptor miniFine-tuned GPT-5 miniGAGeneral-purpose coding and writingRetires 2026-09-01
Kimi K2.7 CodeMoonshot AIGANo task grouping published
Kimi K3Moonshot AIGANo task grouping published
Grok 4.5xAIGANo task grouping published
Grok 4.6xAIGANo task grouping published

Four columns, and every one of them is a published fact. Status distinguishes generally available from public preview. The task grouping is GitHub’s own categorisation from its model-comparison documentation; an empty cell means GitHub lists the model without placing it in a task area, which is not a judgement about quality.

By task area

The same inventory, sliced by what GitHub says each group is for.

General-purpose coding

Model data last verified2026-09-01Model availability changes frequently.

GitHub's description: common development tasks needing a balance of quality, speed and cost efficiency, and a good default when you have no specific requirement.

General-purpose coding models
ModelProviderStatusGitHub groups it under
GPT-5 miniOpenAIGAGeneral-purpose coding and writing; Deep reasoning and debugging; Working with visuals
GPT-5.3-CodexOpenAIGAGeneral-purpose coding and writing
GPT-5.6 TerraOpenAIGAGeneral-purpose coding and writing
MAI-Code-1-FlashMicrosoftGAGeneral-purpose coding and writing; Fast help with simple or repetitive tasksRetires 2026-09-10
Raptor miniFine-tuned GPT-5 miniGAGeneral-purpose coding and writingRetires 2026-09-01

Fast, simple or repetitive work

Model data last verified2026-09-01Model availability changes frequently.

Optimised for speed and responsiveness. GitHub suggests them for quick edits, utility functions, syntax help and lightweight prototyping.

Models optimised for speed
ModelProviderStatusGitHub groups it under
GPT-5.6 LunaOpenAIGAFast help with simple or repetitive tasks
Claude Haiku 4.5AnthropicGAFast help with simple or repetitive tasks
MAI-Code-1-FlashMicrosoftGAGeneral-purpose coding and writing; Fast help with simple or repetitive tasksRetires 2026-09-10

Deep reasoning and debugging

Model data last verified2026-09-01Model availability changes frequently.

For tasks that require step-by-step reasoning, complex decision-making, or high-context awareness.

Reasoning-focused models
ModelProviderStatusGitHub groups it under
GPT-5 miniOpenAIGAGeneral-purpose coding and writing; Deep reasoning and debugging; Working with visuals
GPT-5.5OpenAIGADeep reasoning and debugging
GPT-5.6 SolOpenAIGADeep reasoning and debugging
Claude Opus 4.7AnthropicGADeep reasoning and debugging
Claude Sonnet 4.6AnthropicGADeep reasoning and debugging; Working with visualsRetires 2026-09-01
Gemini 3.1 ProGooglePublic previewDeep reasoning and debugging; Working with visualsRetires 2026-09-01

Visual reasoning

Model data last verified2026-09-01Model availability changes frequently.

For questions about screenshots, diagrams, UI components or other visual input. These models support multimodal input.

Multimodal models
ModelProviderStatusGitHub groups it under
GPT-5 miniOpenAIGAGeneral-purpose coding and writing; Deep reasoning and debugging; Working with visuals
Claude Sonnet 4.6AnthropicGADeep reasoning and debugging; Working with visualsRetires 2026-09-01
Gemini 3.1 ProGooglePublic previewDeep reasoning and debugging; Working with visualsRetires 2026-09-01

Model retirement

Model data last verified2026-09-01Model availability changes frequently.
Models with an announced retirement date
ModelProviderStatusRetires
Claude Opus 4.5AnthropicGA2026-09-01Suggested alternative: Claude Opus 4.7, Claude Opus 4.8 or Claude Opus 5
Claude Opus 4.6AnthropicGA2026-09-01Suggested alternative: Claude Opus 4.7, Claude Opus 4.8 or Claude Opus 5
Claude Sonnet 4.5AnthropicGA2026-09-01Suggested alternative: Claude Sonnet 5
Claude Sonnet 4.6AnthropicGA2026-09-01Suggested alternative: Claude Sonnet 5GitHub states Claude Sonnet 4.6 remains available to individual Copilot subscribers on annual plans, so they retain a Sonnet offering; the deprecation does not apply to those customers.
Gemini 3.1 ProGooglePublic preview2026-09-01Suggested alternative: Gemini 3.6 Flash
MAI-Code-1-FlashMicrosoftGA2026-09-10Suggested alternative: MAI-Code-1.1-FlashAlso the named successor for Raptor mini, which is deprecated on 1 September — nine days before this model is. Anyone following that migration path should go straight to MAI-Code-1.1-Flash.
Raptor miniFine-tuned GPT-5 miniGA2026-09-01Suggested alternative: MAI-Code-1-FlashGitHub names MAI-Code-1-Flash as the replacement, and MAI-Code-1-Flash is itself deprecated on 10 September. MAI-Code-1.1-Flash is the destination that survives both dates.

Retirement is a normal, scheduled part of the product rather than something sudden. Three things about it matter more than the specific dates below.

Successors are named inconsistently. Some retirements point at a replacement; others simply end. A retirement is not always an instruction about what to move to — and where one is named, it is worth checking that the successor is not itself on the list. GitHub names MAI-Code-1-Flash as the replacement for Raptor mini, which is deprecated on 1 September 2026; MAI-Code-1-Flash is deprecated on 10 September 2026. Following the migration literally means moving twice in nine days. MAI-Code-1.1-Flash is the destination that survives both dates.

Exceptions exist. One model above remains available to individual subscribers on annual plans despite the general deprecation, which means “is this model available” can depend on who is asking.

Anything naming a model has a shelf life. A prompt file pinning one, a workflow specifying one, a runbook mentioning one — each becomes an edit on a date you did not choose. This is the strongest practical argument for auto selection: a session on auto continues working through a retirement that would have broken a pinned configuration, with no action from anyone.

What the four columns are, precisely

Being explicit about what each column claims prevents over-reading the table.

Model is the name as GitHub lists it. This is what appears in a model selector and what you would write if pinning one.

Provider is who makes it. Useful for recognising family conventions and for understanding that the inventory spans several suppliers rather than one.

Status is generally available or public preview. This is a statement about stability of commitment rather than about quality — a preview model is not worse, it is less settled.

Task grouping is GitHub’s own categorisation from its model-comparison documentation, reproduced rather than interpreted. Where a cell is empty, GitHub lists the model as supported without recommending it for a particular kind of work.

What is deliberately not a column: context window size, coding score, speed rating, price. The first is published inconsistently and matters less than people expect; the rest are either unpublished or belong on a billing page that changes independently of this one.

Providers, and what each brings

Seven providers is unusual for a single product, and the composition tells you something about how GitHub is positioning Copilot.

OpenAI supplies the largest share of the inventory, spanning the full range from nano-scale fast models to reasoning-focused ones, plus a coding-specialised variant. This is the deepest single-provider bench in the list.

Anthropic supplies the second largest, organised in a clear capability ladder — Haiku for speed, Sonnet for balance, Opus for reasoning — with several Opus versions available simultaneously rather than only the newest. That simultaneity is worth noticing: it means teams that validated against a specific version are not forced forward on GitHub’s schedule.

Google contributes Flash models across several versions plus one Pro model in preview, weighting its presence toward the fast end.

Microsoft contributes coding-specialised Flash models, one of which is retiring in favour of its own successor — an unusually clean retirement, since the replacement is a point-release of the same family.

Moonshot AI and xAI each contribute a small number, broadening the supplier base.

Raptor mini is listed with a provider of “Fine-tuned GPT-5 mini”, which is a different kind of entry: a model adapted for a purpose rather than a foundation model in its own right.

Preview versus generally available

Two statuses appear in the tables and the distinction has practical weight.

Generally available means the model is a supported part of the product. Retirement, when it comes, is announced in advance with a date.

Public preview means it is available to use and its behaviour, availability and terms may change with less notice. Two entries currently carry it: one Anthropic fast-mode variant and one Google Pro model.

Preview is not a warning about quality. It is a statement about stability of commitment. The practical rule: preview models are fine to use and a poor choice to build a dependency on. If a workflow, prompt file or team convention names a preview model, that is a maintenance obligation you have taken on knowingly or otherwise.

How to use a comparison page without being misled

Comparison pages have a characteristic failure: they make a decision feel more determinate than it is. Three habits guard against it.

Check the date before the content. This page carries 2026-09-01 at the top for that reason. A model table without a visible verification date should be treated as unreliable regardless of how authoritative it looks — the information decays quickly and silently.

Read the task grouping, not the name. The durable question is “which category does my task fall into”, and the answer to that outlives every specific model in the category. Names are how you act on the answer; they are not the answer.

Treat absence carefully. A model with no task grouping is not a bad model — GitHub has simply not published a recommendation for it. Reading an empty cell as a negative judgement is a common misreading of tables like this one.

Why there are no scores here

The obvious thing to want on a comparison page is a number per model, and its absence is deliberate.

GitHub does not publish coding-quality scores. What it publishes is what is in the tables above. Anything more specific would have to come from somewhere else, and would then be a different claim wearing GitHub’s authority.

Public benchmarks measure something narrower than they appear to. A score on self-contained programming problems tells you about self-contained programming problems. Your work is a large codebase with local conventions, instruction files and a half-finished migration, and performance on the first does not predict performance on the second as reliably as the number implies.

Copilot is not the raw model. What reaches you has passed through GitHub’s prompting, context assembly and tooling. A standalone benchmark measures a different system.

Numbers age faster than models. A comparison from three months ago may involve one model that has been superseded and another that has retired.

Reading model names as a family tree

32 names is a lot to hold, and there is structure in them that makes the list far more navigable than it first appears.

Version numbers rise within a family and mean nothing across families. GPT-5.6 is newer than GPT-5.4. Claude Opus 4.8 is newer than Claude Opus 4.7. GPT-5.6 and Claude Opus 4.8 have no relationship at all, and treating the numbers as comparable is the most common misreading of a list like this.

The ordering also tells you nothing about what is being retired. Claude Opus 4.5 and 4.6 are both deprecated on 1 September 2026 while 4.7 continues, so “newer” and “supported” are related but not the same question.

Tier names are reliable within a provider. Anthropic’s Haiku, Sonnet and Opus ascend in capability and cost. OpenAI’s nano, mini and unqualified names do the same. Google’s Flash and Pro split fast from capable. Once you know the ladder for a provider, a name you have never seen places itself.

Suffixes carry meaning. “Codex” or “Code” in a name signals coding specialisation. “Fast mode” signals a latency-optimised variant of an otherwise heavyweight model. Note that GitHub does not currently place that variant in any task group, so the suffix tells you more about it than the comparison page does.

Some names encode nothing. Luna, Sol and Terra sit within one family and tell you nothing about relative capability. Where a name is opaque, the task grouping is the only signal available, and that is what the tables are for.

The value of recognising the pattern is not trivia. When a new model appears next month, the pattern tells you roughly where it sits before any documentation has been written about it.

Same family, several versions

One feature of the current inventory deserves comment because it is unusual: a number of families have several versions available at once. Four Claude Opus versions. Several GPT-5 point releases. Three Gemini Flash versions.

This is a deliberate accommodation rather than an accident of housekeeping. Teams that have validated a workflow against a specific version — a prompt file tuned to one model’s behaviour, an evaluation suite, a compliance sign-off — are not forced onto a new version the day it ships.

Two consequences follow.

Newest is not automatically the right choice. If an older version in the same family does what you need, the newer one is a change with a cost and an unquantified benefit.

Old versions still retire. The accommodation is a delay, not an exemption. The retirement table is where that delay ends, and it is the part of this page worth checking on a schedule rather than when curiosity strikes.

What the table cannot tell you

Four things affect which models you can actually use, none of which is in the inventory.

Your plan. Free and Student reach models only through auto selection. Paid plans open manual selection.

Your surface. Availability varies between GitHub.com, the IDEs and the CLI. A model a colleague uses may not appear for you because you are somewhere else.

Your organization’s policy. Administrators can restrict the available set, and auto routes within what is permitted.

Your allowance. Premium requests deduct by a per-model multiplier, so the practical cost of a model depends on how much of your allowance is left. A model that is the obviously right technical choice on the first of the month may be a harder call on the twenty-eighth, which is a genuine consideration rather than a theoretical one for anyone working at volume.

Reading this page well

Do not memorise it. These are names that will look different in three months. The categories outlast the contents, and the categories are what to carry away.

Use it to check a name. The most common real use of a page like this is confirming that a model someone mentioned in an article or a code review still exists, and is not about to retire.

Use the retirement table as a to-do list. It is the only part of this page with a deadline attached.

Prefer auto unless you have a reason. The strongest argument is not routing quality; it is that auto absorbs retirement without you doing anything, and retirement is the one thing on this page guaranteed to happen.

Comparing for yourself

If a comparison table cannot tell you which model suits your work — and it cannot — the question of what does is worth answering rather than leaving hanging.

The only reliable method is to try candidates on work you have already done, where you know the right answer.

Choose a task with a known outcome. A bug you eventually found. A refactor you completed. Something that took real effort, because trivial tasks do not discriminate between models.

Hold everything else constant. Same prompt, same files supplied, same instruction files in force. A comparison where the input varied tells you nothing, and it is easy to vary the input without noticing — a different opening sentence counts.

Judge on what you actually care about. Not whether the answer sounded authoritative. Did it identify the real cause or a plausible-sounding wrong one? Did it stay inside the scope you set? Did it notice a constraint you had not stated? Was it fast enough that the interaction was worth having?

Run each more than once. These systems are non-deterministic. A model that succeeds once and fails twice is not better than one that is consistently adequate, and a single run cannot tell the difference.

Write down what you concluded and when. Your finding has the same shelf life as this page. A note saying “tested on the parser bug, March, X handled the cross-module reasoning and Y did not” is far more useful in six months than a remembered preference with no attached reasoning.

What changes, and how fast

Some sense of the rate of change helps calibrate how much to trust any snapshot, including this one.

The inventory changes monthly at least. New models arrive between retirements, and the list is materially different across a quarter.

Retirements are announced ahead of the date. This is the predictable part, and it is the part worth subscribing to. GitHub’s changelog carries them.

Task groupings move. A model’s recommended use can be revised as GitHub learns more about how it performs, without any change to the model itself.

Status changes in one direction, usually. Preview features tend toward general availability rather than away, though neither is promised.

Plan and policy availability changes independently. A model can remain in the inventory while becoming unavailable to you, because your organization changed a setting.

The reason this page carries a seven-day review interval rather than the site’s usual longer one is that four of those five can invalidate it without any announcement reaching a reader.

Verification and maintenance

Everything above renders from a single data file, which is deliberate. A model list restated in three articles is three lists that will disagree within a month, and the disagreement is invisible until a reader acts on the stale one.

The data carries its own verification date, and the pages built on it display that date rather than burying it — a model table without a date is actively misleading, because a reader has no way to judge how much to trust it.

Model facts on this site are also checked by an editorial audit that flags articles referencing models absent from the data, and retired models presented as current. That catches the specific failure where prose and data drift apart — which is the failure that matters, because a table can be updated while the sentence three paragraphs above it still names a model that no longer exists.

None of this makes the page permanently correct. It makes the page honest about when it was last checked, and it makes a stale claim in prose fail a build rather than reach a reader.

Frequently asked questions

Why does my Copilot show fewer models than this list? Three possible reasons, in order of likelihood: your plan restricts selection, your organization’s administrator has limited the available set, or you are on a surface where fewer are offered. None of them is a fault.

Is the newest model always the best choice? No. Within a family a higher version is generally an improvement, but “improvement” is measured across many tasks and yours may not be one of them. If an older version does what you need, switching is a change with a cost and an unquantified benefit.

What happens when a model I use retires? Requests to it stop working after the announced date. Where GitHub names a successor, that is the intended path; where it does not, the task groupings are how you find an equivalent. Sessions on auto are unaffected, which is a quiet but real argument for auto.

Can I still use a model after its retirement date if I have an annual plan? One current retirement carries exactly that exception, which is why the retirement table records notes rather than only dates. Do not generalise from it — read the specific announcement.

Does model choice affect code completions as well as chat? Model selection applies primarily to chat and agentic interactions. Do not assume inline completions are governed by your selected model.

How do I know which model answered? Some surfaces show it. The CLI displays the model in use, which is useful when running on auto and wanting to know what it chose.

Next

Choosing a model is the decision framework. Models explained covers the concepts — why there are so many, what auto does, how credits work.

For the levers that affect output more than model choice does, context and instructions.

Sources

Every version-sensitive claim on this page was checked against first-party documentation. Only sources actually used are listed.

Primary sources