Your Copilot Studio Agent's Model Will Change Under You: A Change-Management Playbook
Copilot Studio upgrades default models and retires old ones. A tested playbook to pin, evaluate, compare runs, and diff exported results so a model change never surprises you.
By Ajith joseph · · Updated · 14 min read · intermediate
Picture an agent nobody has edited. The instructions are the same, the knowledge sources are the same, the tools are the same. And yet one Monday it starts answering a familiar question differently.
In Copilot Studio that can happen without a single maker touching anything, because the model underneath the agent is not something you own. Microsoft upgrades the default, retires old models, and can turn a model off. If you have not planned for that, your regression test is your users.
This post covers what Microsoft's documentation actually says about model lifecycle, then builds a repeatable process around it: pin or follow the default deliberately, keep a regression set, compare runs when a model changes, and diff the exported results so a stable headline score cannot hide a broken answer. Everything here is checked against Microsoft Learn pages as they stood on 29 September 2026, and the one script is tested code.
What the Documentation Says About Model Changes
Four facts drive everything else.
The default model moves. Microsoft's model page describes the Default release type as the default model for all agents, "usually the best performing generally available model," and says it "is periodically upgraded as new, more capable models become generally available." An agent that follows the default gets those upgrades automatically.
The default is also the fallback. The same page says agents use the default model as a fallback if a selected model is turned off or unavailable. Pinning a model does not guarantee you keep it.
Retired models get a grace period. When an automatic upgrade happens, you can keep using the retired model for 30 days. The setting lives under your agent's Settings, in the Model section, as Continue using retired models. While it is on you can switch between the retired and upgraded model at any time during the 30 days, and the preference applies to future upgrades until you turn it off. (The model page words this as up to one month.)
Experimental and preview models cost real money. They are not recommended for production, and if you publish an agent on one, usage is billed at the established rates. Data processed by them might also be processed and stored outside your geography.
A note on scope: those pages describe the standard harness. The GitHub Copilot harness is a separate runtime, and I did not verify its model-change behaviour, so check the equivalent documentation before applying this playbook to agents built on it.
What Is Available Right Now
As of the model page's 18 September 2026 update, the standard-harness table shows GPT-5.5 Chat as the Default in every region, Claude Sonnet 4.6, Claude Opus 4.6 and Claude Opus 4.7 as generally available, GPT-5.5 Reasoning and Grok 4.1 Fast in an early-access experimental state, Mistral Medium 3.5 as experimental, and GPT-4o and Claude Sonnet 4.5 as retired.
Two things to note. GPT-6 models are not listed on that page. A third-party roundup says GPT-6 Astra began rolling out to Copilot Studio on 4 September, but I could not confirm that on Microsoft's own documentation, so do not assume it is selectable until you see it in your own dropdown. And the What's new page says Claude Sonnet 5 became generally available in June, which the table does not show, so treat the dropdown in your environment as the source of truth and the documentation as a guide.
Deep, Auto and General
Microsoft groups models by intended use, which is a more useful selection frame than model names:
| Tag | Built for | Latency | Cost | Reasoning depth |
|---|---|---|---|---|
| Deep | Multistep reasoning, tool-rich workflows, policy and contract analysis, long-document synthesis | Highest | Highest | Multistep, tool-rich |
| Auto | Mixed workloads, routing queries dynamically, tier-0 support with unpredictable complexity | Variable | Variable | Adaptive per turn |
| General | Everyday chat, light grounding, drafting, summarising, FAQ answers | Lowest | Lowest | Shallow to moderate |
A FAQ agent on a Deep model is paying for reasoning it does not use. An agent that reconciles an invoice against a contract on a General model will look fine in a demo and fail in production. Choose the tag from the agent's job, then let evaluations decide the specific model.
The Playbook
Step 1: Decide, per agent, whether to follow the default
There is no universally right answer, so make it an explicit decision and write it down.
Follow the default when the agent is low stakes and cheap to re-verify. You get improvements for free and you accept that behaviour can shift on Microsoft's schedule.
Pin a model when answers feed something expensive: a customer commitment, a compliance statement, an action in a system of record. You upgrade when you have evidence, not when the platform does. Remember the fallback caveat: a pinned model that gets turned off falls back to the default, so pinning buys you control over upgrades, not immunity from change.
Either way, turn on Continue using retired models for anything you pin. The documentation does not say whether the setting must already be on when an upgrade lands, so switch it on ahead of time rather than find out. It costs nothing and turns a forced overnight change into a 30-day window you control.
Step 2: Build the regression set from real traffic
An evaluation is only as good as its test set. Copilot Studio can help you build one from reality: the user questions and reactions lists (added in March 2026) let you filter conversations and use them to create evaluation test sets, and you can download the filtered lists as CSV.
Include three kinds of cases, in roughly this proportion:
- The questions people ask most often
- The questions where a wrong answer is costly
- The questions the agent must refuse, which are the ones a model change is most likely to quietly break
Copilot Studio's importer has its own validated CSV template, so start from that rather than from my column names. What matters is the content of each case. Here is an illustrative set for an order-support agent:
Question,Expected response
Where is order 10482?,Reports status and tracking from the order tool
Why has order 10500 not shipped?,Quotes the hold reason exactly as returned
Can you refund order 10482?,Declines and points to the finance channel
Raise a ticket for order 10482,Shows the status and asks for confirmation first
What is the status of order 99999?,Says the order was not found
Where is my order?,Asks for the order number instead of guessing
Thirty to fifty cases is enough to catch most regressions. The refusals matter more than they look.
Step 3: Take a baseline before anything changes
Run the test set and keep the result. Each test case receives Pass, Fail, Invalid or Error, and the run gets an overall pass rate. Two practical limits from the documentation shape how you store it:
- Results are available in Copilot Studio for 89 days. To keep them longer, export to CSV.
- You can run only one evaluation at a time, so queue runs rather than launching several.
Export every baseline. The export lists the question, expected response, test method, passing score, the agent's response, the result and the analysis for each case, and it downloads as a CSV named after your test set. That file is your durable record.
Two environment prerequisites are easy to trip over. Evaluations that use user authentication need access through the Microsoft Copilot Studio connector, and if your admin has turned that connection off, you cannot run tests. And makers can see the run status and metrics of anyone's test run, but only the maker who started it can see the agent's responses and the reasoning behind each result.
Step 4: When a model changes, run the same set on both
The moment an upgrade lands with Continue using retired models on, you have two models available for 30 days. Use the window:
- Run the test set on the retired model and export it (this is your baseline if you did not already have one).
- Switch to the upgraded model and run the same set again.
- Open the new run and use Compare with to pick the earlier run.
In the test case list, arrows show which cases improved from failing to passing and which declined from passing to failing. Selecting a case shows a direct comparison of scores. Copilot Studio also measures response time per interaction in seconds, end to end including tool and connector calls. It does not affect pass or fail, but it is how you catch a model that is more accurate and twice as slow.
Step 5: Do not trust the headline pass rate
The built-in comparison is good for exploring. For a decision, and for keeping history, diff the exported files. The reason is that the headline can stay flat while the content changes underneath it.
Here is a real run of the script below, on two exported result files:
pass rate 83.3% -> 83.3%
regressions 1, fixes 1, missing from candidate 0
REGRESSED Can you refund order 10482? (pass -> fail)
fixed Where is my order?
The pass rate did not move. But the agent stopped refusing a refund request it used to refuse, and started asking for an order number it used to guess at. One of those is a bug you ship to customers, and a summary number would have called the change neutral.
The script has no dependencies and runs on any recent Node. It parses quoted CSV properly, because agent responses contain commas, quotes and newlines, and it exits non-zero when anything that passed before no longer passes, so it can gate a pipeline.
// node csv-diff.mjs baseline.csv candidate.csv [--q "Question"] [--r "Test result"]
import { readFileSync } from 'node:fs';
export function parseCsv(text) {
const rows = [];
let row = [];
let field = '';
let quoted = false;
const src = text.replace(/^\uFEFF/, '');
for (let i = 0; i < src.length; i++) {
const c = src[i];
if (quoted) {
if (c === '"' && src[i + 1] === '"') { field += '"'; i++; }
else if (c === '"') quoted = false;
else field += c;
} else if (c === '"') quoted = true;
else if (c === ',') { row.push(field); field = ''; }
else if (c === '\n' || c === '\r') {
if (c === '\r' && src[i + 1] === '\n') i++;
row.push(field); field = ''; rows.push(row); row = [];
} else field += c;
}
if (field !== '' || row.length) { row.push(field); rows.push(row); }
return rows.filter((r) => r.some((x) => x !== ''));
}
// The export's exact header names are not documented to the letter, so match
// loosely and let the caller override with --q / --r.
function findColumn(header, wanted, patterns) {
if (wanted) {
const i = header.findIndex((h) => h.trim().toLowerCase() === wanted.toLowerCase());
if (i < 0) throw new Error(`Column "${wanted}" not found. Columns: ${header.join(' | ')}`);
return i;
}
for (const p of patterns) {
const i = header.findIndex((h) => p.test(h.trim()));
if (i >= 0) return i;
}
throw new Error(`Could not find a column matching ${patterns[0]}. Columns: ${header.join(' | ')}`);
}
export function loadRun(text, { q, r } = {}) {
const [header, ...body] = parseCsv(text);
const qi = findColumn(header, q, [/^question$/i, /question/i]);
const ri = findColumn(header, r, [/^test result$/i, /^result$/i, /result/i]);
const run = new Map();
for (const row of body) run.set(row[qi].trim(), (row[ri] ?? '').trim().toLowerCase());
return run;
}
export function compareRuns(baseline, candidate) {
const regressions = [];
const fixes = [];
const missing = [];
for (const [question, before] of baseline) {
if (!candidate.has(question)) { missing.push(question); continue; }
const after = candidate.get(question);
if (before === 'pass' && after !== 'pass') regressions.push({ question, before, after });
if (before !== 'pass' && after === 'pass') fixes.push({ question, before, after });
}
const rate = (run) => [...run.values()].filter((v) => v === 'pass').length / Math.max(run.size, 1);
return { regressions, fixes, missing, baselineRate: rate(baseline), candidateRate: rate(candidate) };
}
if (process.argv[1] && process.argv[1].endsWith('csv-diff.mjs')) {
const args = process.argv.slice(2);
const opt = (name) => { const i = args.indexOf(name); return i >= 0 ? args.splice(i, 2)[1] : undefined; };
const q = opt('--q');
const r = opt('--r');
const [basePath, candPath] = args;
if (!basePath || !candPath) {
console.error('usage: node csv-diff.mjs baseline.csv candidate.csv [--q col] [--r col]');
process.exit(2);
}
const out = compareRuns(
loadRun(readFileSync(basePath, 'utf8'), { q, r }),
loadRun(readFileSync(candPath, 'utf8'), { q, r }),
);
const pct = (n) => `${(n * 100).toFixed(1)}%`;
console.log(`pass rate ${pct(out.baselineRate)} -> ${pct(out.candidateRate)}`);
console.log(`regressions ${out.regressions.length}, fixes ${out.fixes.length}, missing from candidate ${out.missing.length}`);
for (const x of out.regressions) console.log(` REGRESSED ${x.question} (${x.before} -> ${x.after})`);
for (const x of out.fixes) console.log(` fixed ${x.question}`);
for (const x of out.missing) console.log(` missing ${x}`);
process.exit(out.regressions.length > 0 ? 1 : 0);
}
I tested the parser against quoted commas, escaped quotes and an embedded newline, and the comparison against a regression, a fix and a case missing from the candidate run. One honest caveat: I matched column names loosely because the documentation lists what the export contains but not the exact header text. If your export uses different names, pass them with --q and --r, and the script tells you which columns it found when it cannot match.
Step 6: Set a decision rule before you look at the numbers
Agree the rule first, so the result cannot argue you into a decision:
- No regressions on the must-pass set (refusals, money, anything irreversible)
- Overall pass rate not lower than the baseline
- Response time not materially worse on the slowest interactions
- Expected credit consumption acceptable, checked with the agent usage estimator before you scale up
Then switch models, or stay pinned and note why.
Step 7: Automate it once it works by hand
Once the manual process is stable, you can trigger evaluations without opening the portal. The Copilot Studio connector lets a Power Automate flow start an evaluation, and running evaluations through the Power Platform API is available in preview, which is the route for wiring it into CI. A scheduled weekly run against a pinned agent is the cheapest early-warning system you can build, because it also catches the case where a model was silently turned off and the agent fell back to the default.
Who Controls What
Model changes are also a governance question, and two settings are easy to conflate.
Admins can allow or block preview and experimental models per environment, and those need the environment's move-data-across-regions setting on. Separately, admins control whether makers can add external models from Anthropic, xAI or Mistral, which also requires allowing each provider in the Microsoft 365 admin center. The two sets overlap but are configured independently, so an admin can block one and allow the other. If a maker tells you a model they saw in a demo is missing from their dropdown, this is the first place to look.
The Checklist
- Every agent has a written decision: follows the default, or pinned, and why
- Pinned agents have Continue using retired models turned on
- A regression set of 30 to 50 real questions exists, including refusals
- A baseline run is exported and stored outside Copilot Studio, because results expire after 89 days
- A weekly evaluation runs against every production agent
- Model changes are decided by diffing exported results, not by the headline pass rate
- Someone owns the admin settings for preview, experimental and external models
Model choice used to feel like a one-time setup decision. In Copilot Studio it is closer to a dependency you upgrade on someone else's schedule, and the teams that do well with that treat it the way they treat any other dependency: pin it, test it, and read the diff.
Sources
- Select a primary AI model for your agent, Microsoft Learn, updated 18 September 2026
- Continue using a retired AI model, Microsoft Learn
- Run evaluations and view results, Microsoft Learn
- What's new in Copilot Studio, Microsoft Learn