Agentic AI | Microsoft Copilot Blog http://approjects.co.za/?big=en-us/microsoft-copilot/blog/copilot-studio/cs-topic/agentic-ai/ Tue, 21 Jul 2026 23:56:56 +0000 en-US hourly 1 https://wordpress.org/?v=6.9.4 Extending agentic Microsoft Dynamics 365 Sales with MCP http://approjects.co.za/?big=en-us/microsoft-copilot/blog/copilot-studio/extending-agentic-microsoft-dynamics-365-sales-with-mcp/ Tue, 21 Jul 2026 19:20:00 +0000 http://approjects.co.za/?big=en-us/microsoft-copilot/blog/?post_type=copilot&p=8252 MCP extends Dynamics 365 Sales agents with external data, tools, and actions, helping connect sales processes across business systems.

The post Extending agentic Microsoft Dynamics 365 Sales with MCP appeared first on Microsoft Copilot Blog.

]]>
Open by design: Dynamics 365 Sales and first-party agents are extensible through the Model Context Protocol (MCP)

AI agents are only as capable as the tools, skills, and information they can easily connect to. In sales, the intelligence that drives great decisions—buying signals, conversation insights, market and account context—often lives outside the customer relationship management (CRM). Getting it into the flow of work has traditionally meant costly integrations or data-duplication projects only the largest companies could sustain.

The Model Context Protocol, or MCP, is an open standard that lets AI agents securely connect to external tools, systems, and data sources. MCP helps organizations unlock the full value of their CRM data by connecting agents to the systems, tools, and knowledge they already use. It extends first-party agents, unifies fragmented business data, brings CRM intelligence into any experience, and accelerates the development of custom agents. For Dynamics 365 Sales, that means agents can securely tap trusted intelligence in real time without requiring sellers or IT teams to rebuild integrations or duplicate data. The value shifts from data plumbing to intelligent action, giving sellers and sales leaders a more complete, connected, and actionable experience inside Dynamics 365 Sales.

Our Microsoft-built sales agents in Dynamics 365 Sales—including the Sales Qualification Agent and Sales Opportunity Agent—are built on Microsoft Copilot Studio, where MCP tools plug in natively. You can bring your existing partners, licenses, and trusted data directly into Dynamics 365 first-party agent experiences without standing up new CRM integrations or duplicating data.

The result
Higher-value Dynamics 365 first-party agents that reason over trusted third-party intelligence and improve output and results.
Lower friction to adoption, because you can use the tools you already have and trust across multiple experiences.
Clear differentiation, where the Microsoft experience can better support agentic sales use cases and workflows, connecting CRM data, partner intelligence, and business context across Dynamics 365 Sales, Microsoft 365, and beyond.
MCP helps turn Dynamics 365 into an agent-powered engine for action
Data and sales intelligence are just the beginning. Because MCP is an open protocol, the same model that brings trusted data into your agents can connect them to nearly any external system or capability, from research and enrichment to workflow, analytics, and tools your teams have not even adopted yet. Today’s launch partners are a powerful starting point and as the MCP ecosystem grows, so does what your agents can do. This strengthens the broader Microsoft Copilot ecosystem and reinforces Microsoft’s agent extensibility.

Not just first-party agents
These partner MCP servers also work with any custom agent built in Copilot Studio, not only our first-party Dynamics 365 Sales agents. Customers can bring this same flexibility into the agents they build for their own specific workflows.

If you already subscribe to these solutions, this is immediate added value. Your existing subscription now flows directly into your agentic workflows in Dynamics 365 Sales, putting data and capabilities you already have to work inside the agent experience, with no duplicate integration or added licensing.

Announcing new MCP partners

ZoomInfo
ZoomInfo: Connect ZoomInfo’s verified business-to-business (B2B) data and context to Microsoft Copilot Studio via ZoomInfo’s headless go-to-market (GTM) context layer, GTM.ai. This surfaces 400M contacts, 100M companies, 120M direct dials, and billions of buying signals inside Microsoft Dynamics 365. In Dynamics 365 Sales agents, customers can enrich contact and account records inline with verified contacts, firmographics, technographics, funding history, and intent signals without manual entry or tab switching.

Dun & Bradstreet
Dun & Bradstreet: Anchored by the global standard D-U-N-S® Number identifier, the D&B Commercial Graph MCP server provides the foundational context layer that allows AI agents to understand business identity, relationships, and risk across the global economy, delivering outputs that are consistent, explainable, and auditable. Leverage the Commercial Graph in Copilot Studio to create custom agents and workflows across finance and credit, supply chain, compliance, risk, sales & marketing, and enterprise master data use cases.

LeadIQ
LeadIQ: LeadIQ connects directly into the Sales Qualification Agent, the same AI agent already researching and scoring every lead that enters Dynamics 365 Sales. Instead of bolting on a separate enrichment step, LeadIQ becomes a knowledge source inside that existing research flow: as the agent evaluates a lead, it calls LeadIQ for a verified work email and phone number, and the result lands in the lead’s Deeper insights alongside the agent’s other findings, cited, sourced, and ready before a seller ever opens the record. There’s no new tool to learn and no separate enrichment step to run. Verification just becomes part of your Dynamics 365 workflow.

Draup
Draup: Draup brings context-rich account and market intelligence into the Sales Qualification Agent—hiring surges, funding events, IT spend shifts, executive moves—then converts those buying signals into clear next actions. Every lead arrives researched, prioritized, and paired with a recommended play.

Gong
Gong: Bring Gong’s account and deal insights directly into the Dynamics 365 agent experience. With real-time customer insights like objections raised, risks flagged, or agreed-upon next steps right at sellers’ fingertips, GTM teams can execute on informed next-best actions.

Enlyft
Enlyft: Enlyft helps B2B technology vendors and their partners turn deep account intelligence—technology adoption, buying propensity, and competitive footprint into pipeline. Our MCP connector brings that intelligence into Dynamics 365 Sales agents, so a qualification agent reasons over a prospect’s real tech stack, competitive install base, and fit signals before a seller ever engages.

HG Insights
HG Insights. HG Insights extends Dynamics 365 Sales agents with trusted B2B market context through Model Context Protocol (MCP), giving sellers direct access to company intelligence—such as technology adoption, cloud and AI maturity, IT investment signals, firmographics, and market context—right inside the agent experience. By combining CRM data with this external market intelligence, Dynamics 365 Sales agents help reps prioritize accounts, uncover whitespace opportunities, and recommend more informed next actions throughout the sales cycle. HG Insights is trusted by 95% of Fortune 1000 B2B tech companies and every major hyperscaler to find, target, and convert their best opportunities faster, and this integration brings that same market intelligence to Dynamics 365 users.

These launch partners highlight how powerful and extensible Dynamics 365 Sales agents can become with MCP: more connected, more contextual, and easier to extend with the partner intelligence customers already trust. Customer can use parter MCP tools with Microsoft-built agents such as Sales Qualification Agent and Sales Opportunity Agent, or bring those same tools into custom agents built in Copilot Studio.

To get started, choose the sales scenario you want to improve, connect the right partner MCP tool, and define the insight or action the agent should support. Learn how to add custom research insights and MCP tools to Dynamics 365 Sales agents.

Want to become a partner? The MCP ecosystem on Copilot Studio is open and growing. Certify your MCP server to make your solution available to Dynamics 365 Sales customers.

The post Extending agentic Microsoft Dynamics 365 Sales with MCP appeared first on Microsoft Copilot Blog.

]]>
Building reliable voice agents: A practical guide http://approjects.co.za/?big=en-us/microsoft-copilot/blog/copilot-studio/building-reliable-voice-agents-a-practical-guide/ Thu, 25 Jun 2026 16:00:00 +0000 http://approjects.co.za/?big=en-us/microsoft-copilot/blog/?post_type=copilot&p=8035 Learn how to design AI voice agents that handle real conversations, stay reliable under pressure, and scale from prototype to production.

The post Building reliable voice agents: A practical guide appeared first on Microsoft Copilot Blog.

]]>
There’s no question that customer-facing AI can carry a conversation. The question is: Can you trust it to complete one?

When customers talk to your agent, they expect voice experiences to be fast, natural, and get them the answers they need. Where’s my order? Can I change the delivery address? Why do I see two charges? When looking for support, they don’t care what stack you used. They care that the agent keeps up, stays on track, and knows when to hand off.

This is a practical playbook for designing customer-facing voice agents that are not just capable, but reliable. The principles apply on any platform; voice reliability is a discipline, not a feature. It’s the decision points, patterns, and checklists that move you from a voice agent prototype to something you’d confidently put in front of customers.

In this guide, we’ll explore:

In customer service, capability may capture attention, but reliability is what earns trust at scale. As customer-facing AI takes on more consequential interactions, reliability may well determine whether automation creates value in your organization—or creates risk.

Why voice agents demand a higher standard of reliability

Traditional customer service systems have been judged primarily on whether they route customers correctly. Modern voice agents are increasingly expected to understand intent, access business systems, complete transactions, and recover when conversations go off script. Each new capability expands what customers can accomplish—but also raises the consequences of getting something wrong.

Reliability is harder when agents take action.

Voice agent conversations also feel more “live” than chatbot conversations. Customers interrupt, change their minds mid-sentence, and need the agent to remember what they said two turns ago. And because they’re usually calling for help, voice is judged by outcomes: did the issue get resolved, correctly and efficiently?

So customer-facing voice reliability is less about a single accurate answer and more about end-to-end behavior. The voice agent needs to move a conversation from intent to action to confirmation, with guardrails and graceful escalation when automation is no longer the right path.

Good news: you don’t have to clear that bar the same way every time. First, decide what kind of agent the job calls for.

Voice AI options explained: IVR vs. generative vs. real-time

Not every call needs a cutting-edge agent. Over-engineering is its own kind of unreliability. Most platforms let you build across three broad tiers. Our advice is to match the tier to the scenario, not the hype. In general, you can classify voice agents in three tiers:

TierWhat it isBest forTrade-off
Tier 1: Classic interactive voice recognition (IVR)Deterministic menus and prompts using speech-to-text, text-to-speech, and touch-tone (aka dual-tone multi-frequency or DTMF) inputHigh-volume, structured tasks: balance checks, store hours, simple status lookupsPredictable and low-cost, but rigid—callers follow the path you define
Tier 2: Generative AI voiceA model that understands natural speech and generates responses that are grounded in your business dataConsidered the mainstream sweet spot: order tracking, billing questions, appointment changes in the customer’s own wordsFlexible and natural, but needs grounding and guardrails to stay reliable
Tier 3: Premium generative AI with real-time speech-to-speechNative speech-to-speech capability with very low latency, fluid barge-in, and the most natural turn-takingAdvanced or “luxury” experiences where natural, interruption-friendly conversation is the differentiatorHighest capability and most natural feel; reserve it for where that experience moves the needle

Think of real-time voice as the premium tier. It shines when the conversation itself is part of the brand. But many customer-facing scenarios are well served by Tier 1 or Tier 2. Whichever tier you choose, the bar comes down to one word: reliability.

Reliability: The foundation of customer-facing AI

Natural feel, warm tone, flexibility—these voice agent perks only matter if the agent reliably does the job. An agent that drops context or invents a delivery date isn’t delightful; it’s a liability.

Here’s the definition we’ll use: a reliable voice agent consistently completes the customer’s task, handles interruptions and clarifications without losing context, and escalates smoothly—with full context—when human judgment is required.

How do you know if an agent is reliable? We’ll tell you: the same seven behaviors show up in every reliable agent. If yours does all seven, you’re on the right track.

The 7 things every reliable voice agent does

  1. Keeps a clear task thread across changes in phrasing or order.
  2. Grounds answers in the systems that run the business—not guesses.
  3. Confirms key details (the “receipt”) before any consequential action.
  4. Uses voice-specific affordances (DTMF, barge-in, silence detection) to keep calls moving.
  5. Explains what it’s doing while back end actions run.
  6. Recognizes its boundaries and routes to a human.
  7. Leaves the next human with context, not a blank slate.

What reliability looks like in live voice conversations

Here’s an example of each from a real call with an agent from a hypothetical clothing retailer.

1. Keeps a clear task thread.
“Where’s my order—wait, why was I charged twice?” The agent parks the order question, fixes the billing one, then circles back: “That duplicate charge is reversed—now, order #18372 is out for delivery today.”

2. Grounds answers in real systems.
Instead of guessing “three to five days,” the agent reads the live record: “Out for delivery, arriving by 6 PM today.”

3. Confirms the receipt before acting.
Before refunding: “To confirm—cancel the blue jacket on #18372 and refund $89 to your Visa ending 4412—shall I go ahead?” The customer catches a wrong card or item before money moves.

4. Uses voice-specific affordances.
On a noisy line: “I’m having trouble hearing you—type your six-digit order number on your keypad.” Barge-in lets impatient callers cut in; silence detection re-prompts instead of leaving dead air.

5. Explains what it’s doing.
Silence reads as a dropped call, so it narrates: “Give me a moment while I pull up your account—about ten seconds.”

6. Recognizes its boundaries.
“My package never arrived and I want a refund” trips a defined boundary, so it escalates rather than improvising a policy it doesn’t own.

7. Hands off with context.
On transfer it passes a summary: “Identity verified, #18372 marked lost, customer wants a refund”—so the rep picks up mid-stride.

That’s the what of reliable voice agents. Next, the who—because the job of making an agent accurate and trustworthy is almost never owned by just one person.

Who is responsible for voice AI reliability?

Reliability isn’t created by a single feature or team. It emerges from a series of decisions across customer experience, operations, integrations, and governance. Different teams own different parts of that equation, but each contributes to the same outcome: a customer experience that consistently delivers results. Start by identifying which part of reliability you own.

If you ownYour primary goalTypical voice scenariosWhere reliability lives for you
Customer service and support opsDeflect common requestsOrder status, billing questions, appointment schedulingEscalation pathways and consistent outcomes
Contact-center workflowsImprove handle timeIntent triage, case creation, transfer to humanHandoff continuity and edge-case handling
Digital channelsExtend existing chat flowsReschedule, update address, subscription changesContext retention across turns
Systems and platform integrationIntegrate systems safelyAccount lookup, eligibility checks, authenticated actionsData grounding and governance
Custom development and orchestrationCustom user experience (UX) and orchestrationIn-app support, complex multi-step tasksLatency management and tool reliability

You don’t need every piece covered to start—just name the hat you’re wearing today. And now that you have an initial who, let’s move on to how. How do you actually create reliable voice agents?

How to design voice agents around real use cases

Start a voice project by listing features and you’ll get an agent that demos well but struggles in real use. Better: start with a few high-volume scenarios and design around the natural shape of each conversation.

The map below is a starting point. Each scenario needs a primary task, the data to complete it, and an escalation trigger, because nothing is 100% automatable.

Customer scenarioPrimary taskData the agent needsEscalation trigger (example)
Appointment schedulingBook or modify an appointmentAvailability and customer recordNo matching slot/conflict
Order trackingRetrieve delivery statusOrder system and shipping updatesLost package/exception
Billing and paymentsExplain a charge or payment statusInvoice and payment historyDispute, refund request
Service start or stopChange a start date or service optionEligibility and service rulesEligibility failure/safety exception
Account updatesUpdate contact info or preferencesCustomer profileIdentity verification needed

Take order tracking: the task is narrow (“retrieve delivery status”), the data is your order and shipping systems, the trigger is a lost package. Build that end to end before adding billing or returns. One rock-solid scenario beats five shaky ones.

Then build reliability in from the start. Just as every stage of a house build—from the foundation to the framing to the roof—contributes to its strength and stability, every stage of your agent build should contribute to its accuracy and consistency.

A five-pass framework for building reliability into a voice agent

Here’s how to layer reliability in pass by pass, not as a bolt-on at the end.

Pass 1: Define the task and the boundaries

Pick one scenario and write a plain, natural-language success statement: “The customer can check their order status and get an ETA.” Then a boundary statement: “If the order is lost or the customer wants a refund, we hand off to a live rep.”

Those two sentences stop scope creep and give a clean, testable escalation rule. Keep boundaries tight—three or four triggers, not a policy manual.

Pass 2: Design the conversation as a sequence of receipts

Customers can’t see what the agent “stored” unless the agents says it back. Reliable agents use receipts in the form of short confirmations at key points: “Got it—order 18372, shipping to Detroit, latest delivery estimate.” These help head off misunderstandings and interruptions. Issue one whenever the agent captures a key value, and again before any irreversible action.

Pass 3: Use voice-specific controls to keep calls moving

Speech and DTMF input, silence detection and timeouts, latency messages, barge-in, Speech Synthesis Markup Language (SSML), and call transfer aren’t “legacy” capabilities. They’re reliability measures. They help customers recover from recognition errors, give the agent a safe fallback, and prevent dead air.

Pass 4: Ground answers in the systems that matter

Reliability collapses the moment an agent hallucinates an operational fact (delivery window, balance, open slot, etc.). Ask an ungrounded agent when an order will arrive and it might confidently answer, “Thursday.” If that’s wrong, a simple status check becomes a trust problem.

Operational facts should come from systems of record, not model reasoning. And because voice interactions introduce their own opportunities for error, key inputs should be captured carefully: ask once, repeat back, and confirm before taking action.

Pass 5: Prove it works with evaluation-by-scenario

Reliability is demonstrated, not asserted. Build a small per-scenario test set—a dozen realistic calls, including the messy ones (interruptions, wrong inputs, the lost-package path)—and run it whenever you change prompts or integrations. The goal isn’t day-one perfection; it’s catching regressions before customers do.

Together, those passes make up a reusable checklist:

Business-to-consumer (B2C) voice scenario design checklist

  • Scenario is clearly named and outcome-based (not feature-based).
  • Primary task is explicit, plus at least one escalation trigger.
  • Key inputs are captured in a voice-friendly way (ask once, repeat back, confirm).
  • At least one fallback path exists (DTMF option, re-prompt, or transfer).
  • Agent provides “receipts” at key moments so customers can correct course.
  • Long-running actions have a “still working” message to avoid dead air.
  • Handoff includes a short context package for the human.

From prototype to production: What changes?

A prototype can feel great in a demo. Production is different. In a demo, an agent only needs to successfully complete a scenario once. But things change at go-live.

Thousands of customer conversations and edge cases test the agent’s abilities. Customers phrase things differently than your test prompts. They interrupt. They change topics. They provide incomplete information. Your script says “check the status of order #1258;” a real caller says “uh, where’s my stuff?”

The table below provides a simple maturity model for thinking about that progression:

StageWhat you focus onWhat “reliable” means here
PrototypeOne scenario, happy pathConversation is coherent end-to-end
PilotMultiple phrasings and interruptionsAgent recovers from clarifications
ProductionReal data and action-takingGrounded answers and safe actions
ScaleMore scenarios and channelsConsistent behavior and handoff
OptimizeContinuous monitoringQuality improves without regressions

Another important production consideration is where. Through what channel will customers actually engage with the agent? Whether the primary channel is a website, mobile app, or contact-center entry point, choosing that channel early helps you design for that surface’s realities: authentication, user interface (UI) constraints, formatting, and escalation. Picking the primary channel up front can prevent costly rework later.

The why: Earning the right to scale

We’ve covered the what, who, how, and where of reliable voice agents. The final question is why. Why is it so important for organizations to invest in getting this right?

Organizations don’t want AI to be a fun experiment anymore. Like any business asset, voice agents need to deliver value. And reliability is what separates an interesting pilot from a program an organization can confidently scale.

Many organizations can get a voice agent working for a handful of carefully chosen scenarios. The real challenges emerge when they expand: more customers, more channels, and more consequential interactions. That’s when gaps in grounding, escalation, evaluation, and ownership pop up.

A voice agent that loses context, misunderstands requests, or provides incorrect information doesn’t just fail a conversation—it erodes confidence in the broader customer experience. And without customer trust, the opportunity to scale quickly disappears.

The organizations realizing the most value from AI aren’t distinguished by the number of agents they’ve deployed. They’re distinguished by the rigor behind them. Reliability creates the foundation for trust, turning isolated successes into repeatable, governable, and continuously improving customer experiences.

That’s ultimately what this guide has been about: not just how to build a voice agent, but how to build the operational foundation for customer-facing AI.

Building production-ready voice agents with Copilot Studio

The principles in this guide are platform-agnostic by design, and outcomes of course depend on implementation, data, and configuration—but they need a place to come together. Copilot Studio brings together the capabilities designed to help you build reliable customer-facing voice experiences in one platform—from classic IVR through to real-time voice—allowing you to start simple and grow.

The same patterns you’ve seen throughout this guide can be implemented directly in Copilot Studio. Teams can connect agents to systems of record for grounded answers, use voice-specific controls such as DTMF and barge-in to improve call flow, define escalation paths for complex situations, and evaluate agent behavior before deploying changes broadly.

Perhaps most importantly, organizations can start small. A single high-volume scenario—order tracking, appointment scheduling, account updates—can become the foundation for a broader voice strategy. As needs evolve, teams can expand to additional scenarios, channels, and capabilities without rebuilding from scratch.

Ready to get started? Pick one scenario, connect the minimum data required to complete it successfully, and test it end to end in. The most effective voice agents aren’t built all at once—they’re built one reliable customer experience at a time.

Build with confidence

Create reliable voice agents that stay grounded, handle real scenarios, and scale from prototype to production.

Two people review content on a laptop while standing in a shared indoor workspace.

The post Building reliable voice agents: A practical guide appeared first on Microsoft Copilot Blog.

]]>
Who evaluates the evaluators? The data science behind agent evals http://approjects.co.za/?big=en-us/microsoft-copilot/blog/copilot-studio/who-evaluates-the-evaluators-the-data-science-behind-agent-evals/ Thu, 11 Jun 2026 16:00:00 +0000 http://approjects.co.za/?big=en-us/microsoft-copilot/blog/?post_type=copilot&p=7917 An inside look at the data science and evaluation systems helping teams improve agent quality at scale in Copilot Studio.

The post Who evaluates the evaluators? The data science behind agent evals appeared first on Microsoft Copilot Blog.

]]>
At Microsoft Copilot Studio, we talk a lot about ways to make agents better. You can add knowledge sources, fine-tune models, insert prompts, and more. But how do you know whether an agent is actually getting better? How do you detect regressions before your users do? And perhaps most importantly, how do you trust the signals you’re using to make those decisions?

As data scientists, we don’t ship a model without evaluating it. We evaluate it before the first release, and we evaluate it after every meaningful change.

We validate offline, track metrics over time, compare variants, look for regressions, and ask a simple question: Did this change actually make the model better?

When we started building evaluation features in Copilot Studio, we treated them the same way. We asked: Is the evaluation giving the right answer to “is it better”? This is the quality question.

To answer it, we’ll explore three core areas of AI evaluation quality:

  1. The data behind evaluation, and why generated datasets play such an important role in agent testing.
  2. The evaluators (graders) themselves, and how we validate that graders produce reliable signals.
  3. The metrics we use to determine whether those signals are trustworthy enough to support real-world decisions.

Because as organizations increasingly rely on evaluation to improve AI systems, confidence in the evaluation process becomes just as important as confidence in the agent.

From model evaluation to agent evaluation

Traditional machine learning evaluation is relatively well‑defined. You have:

  • Labeled data
  • A clear task
  • A small set of metrics
  • A mostly static input–output mapping

AI agents challenge many of these assumptions.

Agents operate over multi‑turn conversations, adapt to user behavior, use tools, and use implicit reasoning. Accordingly, they are judged across multiple quality dimensions, correctness, completeness, clarity, coherence, tone, and more.

So Copilot Studio ships evaluation features to help makers answer questions like:

  • Did my agent regress after this change?
  • Does quality degrade in longer conversations?
  • Can I trust my agent to behave as expected?

But that immediately raises a second‑order question: How do we know the evaluation features themselves are giving the right answer?

Building the right evaluation data for agents

In data science, we know that evaluation quality is bounded by evaluation data.

If the data you evaluate on is narrow, unrealistic, or biased, the metrics will look confident but be wrong. That principle applies just as much when we evaluate evaluation features as when makers evaluate their agents.

Real data: Grounded, but limited

For makers, Copilot Studio supports importing real production data to evaluate agent performance. Internally, however, we do not use customer data in any form when validating evaluation features.

Instead, we rely on curated examples and generated datasets that allow controlled and systematic testing—what we call semi‑real examples. These include scenarios inspired by conversations shared by design partners, as well as examples curated during feedback cycles.

For us, and for many of our makers, these alone are not sufficient.

Generated data: Scalable, targeted, and intentional

Real and semi‑real examples provide valuable grounding, but in practice, most evaluation workflows (both internally and for makers) rely primarily on generated data. And that isn’t a compromise; it’s an intentional design choice.

Generated data allows evaluation to start earlier, scale faster, and cover a broader range of agent behaviors.

From our perspective as a data science team, generated datasets are essential for supporting the wide variety of agents built in Copilot Studio. They allow us to validate evaluation features across different agent types, domains, and interaction patterns, and to do so at a scale that would not be feasible with curated examples alone.

For makers, the motivations are equally practical:

  • Evaluating before publishing: Generated datasets make it possible to assess agent behavior and quality before the agent is exposed to real users.
  • Limited or restricted access to production data: In many cases, makers do not have access to their agent’s production conversations at all, due to compliance, governance, or organizational policies.
  • Working with production data selectively: Even when production data exists, it often needs filtering or augmentation to support systematic evaluation.

Perhaps the strongest motivation is how easy it is to generate high‑quality evaluation datasets. Copilot Studio enables makers to create test sets that are targeted, repeatable, and aligned with their agent’s intended behavior—without requiring manual data collection.

Evaluating data generation for agent evaluations

Evaluation datasets can be generated in multiple ways. We support multiple data generation strategies because they surface different aspects of agent behavior. When applied together, they give makers practical, high‑coverage evaluation datasets.

Data generation strategies

There are four main types of dataset generation:

  1. Single‑turn generation allows makers to test specific behaviors in isolation. These datasets are easier to reason about and are well‑suited for validating correctness, relevance, and instruction adherence.
  2. Multi‑turn generation adds the complexity of context tracking and conversational dependencies. This is particularly useful for makers building predefined flows or agents whose behavior depends on conversation state.
  3. Knowledge‑based generation tends to produce very concrete, sometimes highly specific questions. These queries are effective for testing grounding and answerability against the agent’s connected knowledge sources.
  4. Topic‑based and instruction‑based generation often lead to more general or exploratory questions. These datasets are useful for identifying unsupported or weakly supported areas—reasonable questions users may ask that fall outside the agent’s main flows.

By combining these generation types, makers can build large and diverse evaluation sets that cover both expected and unexpected usage patterns.

How we evaluate generated queries

Because data generation itself is an evaluation feature, we explicitly assess the quality of generated queries. We use an LLM‑as‑a‑judge methodology to assess dataset quality along several dimensions, including:

  • Relevance. How well queries align with the agent’s intended scope.
  • Interaction naturalness. Whether queries resemble plausible user goals, confusion, and follow‑ups.
  • Human likeness. The extent to which generated queries resemble questions a human would naturally ask.
  • Redundancy. Whether examples add new coverage rather than repeating similar patterns.
  • Intent diversity. The range of user intents represented in queries (for example, informational, troubleshooting, or exploratory).

In addition, we apply generation‑specific measures where appropriate, such as topic coverage for topic‑based generation or grounding for knowledge‑based generation. These assess, for example, whether questions are answerable using the provided sources.

These metrics allow us to reason systematically about whether a generation capability produces datasets that are broad, targeted, and useful for evaluation—without relying on subjective impressions.

Evaluating graders: Assessing the quality of our evaluators

Graders are the evaluators we build to help makers assess their agents. They produce the scores and labels that makers use to understand what works well and what needs to be improved. For that reason, we assess graders explicitly and independently before they are exposed to makers.

What we expect from a high‑quality grader

We treat graders as a system that estimates quality rather than produces one absolute answer. We assess these graders using the same principles we would apply to any automated evaluation system.

Concretely, we ask whether a grader:

  • Measures the intended dimension and only that dimension.
  • Distinguishes between meaningful differences in responses.
  • Behaves consistently across similar inputs.
  • Produces interpretable and stable signals that can support downstream decisions.

A grader that produces reasonable explanations, but inconsistent judgments doesn’t meet the bar.

Purpose‑built datasets for grader assessment

To assess a grader’s quality, we build purpose-built datasets, each tailored to the specific behavior or quality dimension the grader is designed to measure.

Each grader requires targeted datasets designed to measure the specific behavior being evaluated. As a result, the datasets we use for grader evaluation are intentionally designed for that purpose.

In practice, the composition of grader‑specific datasets depends on the grader. For some graders, we rely primarily on human-labeled data. For others, generated data plays a central role, allowing us to construct targeted test cases with known ground truth. Most often, we use a combination of the two, balancing human judgment with scale and control.

A controlled generation methodology

For many graders, we use controlled synthetic datasets with known ground truth.

The process works as follows:

  1. Define a test agent. We start with a well-scoped agent configuration that represents the behavior domain the grader is intended to evaluate.
  2. Generate high-quality queries. Using our data generation capabilities, we create a set of realistic, high-quality user queries aligned with the agent’s scope.
  3. Generate high-quality responses. For each query, we generate responses that meet the expected quality bar for the dimension under evaluation.
  4. Introduce controlled degradations. We then intentionally degrade a subset of these responses in a controlled and traceable way. Each degradation targets a specific failure mode and we explicitly track whether a response was damaged and how.
  5. Use the dataset for evaluation. The resulting dataset contains both intact and intentionally degraded responses, with known ground truth about their quality.

Because we control the transformation applied to each response, we can treat this dataset as labeled. We know which responses should be flagged by the grader, and for what reason.

Measuring grader performance

Now that we know which responses were intentionally degraded and how, we can evaluate graders in a concrete and measurable way. Rather than relying on subjective inspection, we can treat grader assessment as a standard classification problem with known ground truth.

The main metrics we track and optimize when developing graders are true positive rate (TPR) and true negative rate (TNR).

  • TPR measures how often the grader correctly identifies responses that should be flagged. In our context, this reflects the grader’s ability to detect intentionally damaged or low-quality responses when a problem is present.
  • TNR measures how often the grader correctly accepts responses that should not be flagged. This reflects the grader’s ability to avoid false alarms and not penalize responses that meet the expected quality bar.

These metrics capture the core tradeoff every grader must manage: being sensitive enough to catch real issues, while remaining precise enough to avoid over‑penalizing valid responses.

By evaluating graders against datasets with controlled degradations, we can measure TPR and TNR directly, analyze failure modes, and iterate systematically. This allows us to tune grader behavior intentionally—understanding where a grader is too permissive, where it’s too strict, and how changes affect its decision boundaries.

All together, these techniques allow us to move beyond evaluating individual grader performance and toward a broader goal: building evaluation systems whose behavior can be understood, measured, and improved over time.

Bringing rigor to agent evaluation features

So, who evaluates the evaluators? At Copilot Studio, we approach evals with the same rigor we apply to models themselves. Because as teams increasingly rely on evaluation to guide real-world decisions, they need to trust the systems producing those signals.

In this post, we described how we approach that challenge in Copilot Studio: constructing targeted datasets for graders, using controlled generation to create reliable ground truth, and measuring decision accuracy through metrics such as TPR and TNR. These practices help us understand not only whether an evaluation feature works, but how it behaves, where its limitations are, and how it can be improved over time.

Feedback from design partners and customers plays an important role in this process. When real-world examples reveal gaps in a grader or generated dataset, we incorporate those learnings back into our evaluation process to continuously improve the system.

As the industry continues moving from experimental AI systems to production-scale agents, evaluation will become a foundational capability. As AI agents move into production environments, organizations need to trust not just the agents themselves, but the systems used to evaluate them.

For us, rigorous evaluation is a core part of helping teams build and improve agents with confidence. Because better agent decisions start with trustworthy evaluation.

Build more reliable agents

See how generated datasets, grader validation, and scalable testing help improve agent evaluation quality.

Two IT pros in a modern office space gather around a computer that's out of frame.

The post Who evaluates the evaluators? The data science behind agent evals appeared first on Microsoft Copilot Blog.

]]>
Frontier Tuning: Teaching AI to work the way you do http://approjects.co.za/?big=en-us/microsoft-copilot/blog/copilot-studio/frontier-tuning-teaching-ai-to-work-the-way-you-do/ Tue, 02 Jun 2026 17:00:00 +0000 http://approjects.co.za/?big=en-us/microsoft-copilot/blog/?post_type=copilot&p=7902 Frontier Tuning introduces a new way to tune AI agents around your organization’s workflows, internal knowledge, and compliance safeguards.

The post Frontier Tuning: Teaching AI to work the way you do appeared first on Microsoft Copilot Blog.

]]>
Today at Microsoft Build, we introduced Frontier Tuning, a new approach to making AI work the way your business does by applying reinforcement learning inside your compliance boundary with your own data, processes, and conventions. We’re announcing private preview, available through Forward Deployed Engineers, and upcoming availability in Microsoft Copilot Studio and Microsoft Foundry.

Inside Frontier Tuning

Frontier Tuning has three parts that work together: the environment where learning happens; the unique inputs you provide from your own business; and the tuned output models, skills, and harness that the system produces.

A continuously evolving environment. Tuning runs in a managed Reinforcement Learning Environment (RLE) used both for post-training and inference. During training, the system learns from real workflows, tool usage, and eval signals without affecting production systems. At inference it explores multiple frontier and fine-tuned models, from Microsoft AI and OpenAI, across turns to find stronger candidate paths before returning an answer. The system improves continuously as it learns from each interaction.

Your company’s data, domain knowledge, and workflows, in one platform. You bring your business data and know-how into the RLE: content, processes, conventions, terminology, and workflows that collectively define how your business runs. The experience is built to be easy to use, with no need for a data science degree. With a simple, guided approach, teams can bring data in and start tuning right away, enabling more people within your organization to capture the power of tuning.

Tuned models, skills, and harness that stay within your compliance boundary. This system produces tuned models, embeddings, skills, orchestration logic, and a runtime harness. All of this runs on your data, with your controls, without leaving your compliance boundary. The models inherit your access controls, so only people who could see the underlying data can access models built from it. The tools are virtualized, so agents can improve without affecting production systems.

Together, these three pieces form a loop that gets sharper with every agent interaction. As knowledge in your environment grows, the models and harness evolve with it, so your agents keep getting better at the work you actually do.

Fits the way you operate

Frontier Tuning fits into how you already build and operate agents. Users interact with agents tuned on your company’s data and workflows. Makers and developers build and refine these agents in the tools you’re already using, including Microsoft Copilot Studio and Microsoft Foundry.

For example, soon within Copilot Studio, you will be able to access the RLE and use data like transcripts, knowledge bases and Microsoft 365 artifacts to improve the agents you already rely on. We’re also bringing capabilities to Foundry to allow developers to tune agents alongside the tools you already use. Here too, you’ll be able to set up an RLE, bring in your data, and tune models, including Microsoft AI models, and runtime behavior. More details on Foundry support will come in the coming months. And today, Frontier Tuning is available in Private Preview through our Forward Deployed Engineering (FDE) team. FDEs can partner with you end to end: defining the scenario, setting eval criteria, running the tuning process and delivering the agent, all within your environment.

Whether you build in Copilot Studio, develop in Foundry, or partner with an FDE, Frontier Tuning fits how your business already operates.

Frontier Tuning in action

Frontier Tuning is already in the hands of customers. We’ve partnered with a focused set of organizations including Land O’Lakes, EY, Bristol Myers Squibb, Pearson, McKinsey, McCarthy Tétrault, and the Josh Bersin Company.

“Microsoft Frontier Tuning enabled us to generate significantly better Copilot outputs for Communication Coach. The results were more closely aligned with Pearson’s learning science, giving learners clearer, more actionable feedback on how to strengthen workplace communication,” said Gian Paolo Perrucci, Product & Technology Officer, Pearson.

“Microsoft Frontier Tuning is set to transform the tax practice across the global EY organization. By combining a tax-domain–tuned reasoning LLM with our extensive enterprise knowledge and insights from our Tax Advisors, EY is elevating the delivery of tax services. Leveraging client context in Microsoft Work IQ and deep EY expertise, we are tuning an advisory agent within the RLE that will be deployed to 75,000 tax professionals globally in the coming period.”

– Ben Ambrosino, EY Global Tax CTO, EY

“Microsoft Frontier Tuning has given us a powerful way to bring Galileo’s research-backed HR intelligence into bespoke agents inside the Copilot experience. It is one of the most compelling capabilities we have seen for putting deep domain expertise into the daily flow of work.”

– Josh Bersin, Founder and CEO, The Josh Bersin Company

“With Frontier Tuning, we’re teaching the system how Microsoft HR works – capturing organizational knowledge in one connected environment that learns and improves with every use. We partnered with our product teams until the results were undeniable; successful task completion increased from 13% to 87%. Now we’re expanding to more HR workflows.”

– Nathalie D’Hers, CVP Employee Experience

The pattern across these engagements is consistent: when you teach the system how your organization actually works, you get much higher fidelity output and more predictable execution.

Get started

If you’re interested in Frontier Tuning, visit aka.ms/frontiertuning to learn more. We will follow up with additional details.

The post Frontier Tuning: Teaching AI to work the way you do appeared first on Microsoft Copilot Blog.

]]>
Mistral joins Copilot Studio’s growing lineup of model providers http://approjects.co.za/?big=en-us/microsoft-copilot/blog/copilot-studio/mistral-joins-copilot-studios-growing-lineup-of-model-providers/ Thu, 28 May 2026 08:45:00 +0000 Copilot Studio adds Mistral Medium 3.5, expanding model choice with in‑region data control, strong multilingual performance, and admin governance.

The post Mistral joins Copilot Studio’s growing lineup of model providers appeared first on Microsoft Copilot Blog.

]]>
As organizations around the world continue to bring more agents into production, there’s a growing demand for AI models that align with regional expectations around data handling and control. Microsoft Copilot Studio aims to meet this need by combining model flexibility with enterprise-grade governance, empowering teams to choose the best model for a given scenario while maintaining control over how and where data is processed.

Today, we’re expanding that choice with the addition of Mistral Medium 3.5 for agent building and orchestration. Medium 3.5 is currently available worldwide for customers in early release environments.

For organizations in the European Union (EU), this model has the added benefit of providing more flexibility in how their agents are powered, while keeping data processing in-region. It also helps these teams avoid the extra procurement overhead that can come with adding a new model provider.

Medium 3.5, per Mistral, was built for “long-horizon tasks, calling multiple tools reliably, and producing structured output that downstream code can consume.” Reasoning effort is configurable per request, so the same model can answer a quick chat reply—or, alternatively, work through a complex agentic run. You can read more about this model on Mistral’s blog.

Model access and admin controls

We’re rolling out Medium 3.5 for customers in early release environments. As an experimental model, we recommend using it in non-production scenarios while testing and evaluations are completed.

As with all external model providers in Copilot Studio, admins stay in control:

  1. Opt in via the Microsoft 365 admin center to allow the Mistral Medium 3.5 preview for your tenant.
  2. Enable external model providers in the Microsoft Power Platform admin center so makers in your selected environments can select Mistral Medium 3.5 in the Copilot Studio model picker.

Until both switches are on, the model will not appear to makers. This gives IT teams a structured path to pilot, evaluate, and expand usage on their terms.

Get started with Mistral Medium 3.5 in Copilot Studio

  • Admins: Review the enablement guide and model provider terms, then complete the two-step opt-in.
  • Makers: Once your admin has enabled access, open any agent in Copilot Studio and select Mistral Medium 3.5 (Experimental) from the model selector.

Bringing choice and control together

With Mistral Medium 3.5, Copilot Studio continues to expand the range of models available for agent development—while keeping orchestration, governance, and lifecycle management in one place.

The result: customers can choose the right model for each scenario, meet regional and compliance requirements, and scale agents confidently within a unified platform.

More choice. More control. One platform.

Model choice, on your terms

Build and scale agents with flexible models—while keeping control of your data and governance.

A person working on a computer in an open office setting.

The post Mistral joins Copilot Studio’s growing lineup of model providers appeared first on Microsoft Copilot Blog.

]]>
New and improved: Computer-using agents, a new workflows experience, and real-time voice experiences http://approjects.co.za/?big=en-us/microsoft-copilot/blog/copilot-studio/new-and-improved-computer-using-agents-a-new-workflows-experience-and-real-time-voice-experiences/ Tue, 26 May 2026 16:00:00 +0000 Learn what's new in Copilot Studio, May 2026: computer-using agents are now generally available, plus redesigned workflows and Work IQ extensibility.

The post New and improved: Computer-using agents, a new workflows experience, and real-time voice experiences appeared first on Microsoft Copilot Blog.

]]>

Expectations for agents are changing quickly. Teams want to move beyond conversational experiences to systems that can help get work done by interacting with applications, executing workflows, collaborating across tools, and supporting customers more naturally across channels.

But getting there can be tough. Many organizations are still balancing modern, real-time AI experiences with older systems and processes that were never designed to work together. That can leave teams stuck maintaining brittle automations, disconnected processes, and rigid customer interactions that are difficult to evolve. New updates in Microsoft Copilot Studio focus on helping organizations achieve more connected, adaptive automation systems—structured where needed and adaptive where valuable.

Computer use and workflows: Adapt automation to your team’s real work

Traditional automation works best in predictable environments. But many real business processes are anything but predictable. Interfaces change. Vendor portals update unexpectedly. Legacy systems lack APIs entirely. As a result, even relatively simple processes can require constant maintenance just to keep automations running reliably.

That’s the kind of gap computer-using agents are designed to help close.

Computer-using agents are now generally available

With computer-using agents now generally available in Copilot Studio, organizations can build agents that interact directly with websites and desktop applications through the user interface (UI). This helps you automate processes that previously relied on brittle scripts or manual workarounds because the underlying systems lacked APIs.

With the new release also comes new enterprise-ready capabilities designed to help you operationalize UI automation more confidently. Organizations can now manage credentials more securely, choose models best suited for different automation scenarios, and build more resilient automations that can adapt to changing interfaces instead of breaking whenever a screen or webpage changes.

In addition, organizations can now embed computer-using agents directly into multi-step workflows. This feature, now moving into preview, further helps teams combine API-based actions, approvals, business logic, and adaptive UI interactions within the same automation system.

But adaptability automation still needs structure. As organizations scale beyond isolated automations, teams need a way to orchestrate all these in a way that’s easier to understand, maintain, and evolve over time.

That’s where the new workflows experience in Copilot Studio comes in.

A simpler, more intuitive way to build powerful workflows

Now available in early release environments, the redesigned workflows experience introduces a more intuitive visual designer for building and orchestrating agentic automation in one place. Instead of stitching together disconnected tools and logic across multiple surfaces, you can design workflows end-to-end on a unified canvas. This helps you more clearly see how actions, decisions, and AI-powered steps work together across a business process.

A core component of the new experience is the ability to add existing agents directly into workflows. These agent nodes allow you to create automated solutions that keep the scalable reliability of workflows while bringing in AI intelligence when you need it. For example, when a workflow hits a decision that can’t be captured in simple if-then logic—where it needs to use reasoning over context, orchestrate tools, or retrieve knowledge from multiple sources—an agent node can help bridge the gap and make your workflow more effective.

Invoice classification example in Copilot Studio.

The new designer also helps reduce friction for teams building and maintaining these systems. Inline configuration, simplified building blocks, and node-level testing help validate workflow behavior earlier and iterate more quickly. In addition to agent nodes, AI-powered actions like classification, content generation, and decision support can now be incorporated directly into the workflow.

Together, these updates help organizations combine deterministic orchestration with adaptive execution—structured where needed, adaptive where valuable.

How Graebel combines flows and computer-using agents

Graebel, a global leader in talent mobility, processes thousands of employee relocation requests each year. Most of these come in as unstructured emails filled with unique instructions, attachments, and edge cases. Because Graebel’s proprietary Global Connect platform lacked API support, earlier automation efforts proved too rigid to keep up with the variability of real-world requests. That meant lots of manual input and handoffs—which automation was supposed to solve. Basically, Graebel needed automation that could use reasoning, not just click.

Working with GET AI and Microsoft, Graebel built the Graebel Service Order Agent in Copilot Studio using computer use capabilities to help automate the process end to end. The agent can interpret incoming emails, validate requests against business rules, operate Global Connect directly through the UI, and escalate exceptions through workflows when needed.

By adopting Microsoft Copilot Studio and AI agents, we’ve moved beyond traditional automation to a more intelligent, scalable operating model. This initiative strengthens our ability to serve clients faster and more accurately while positioning Graebel for long-term growth.

—Matt Brownlee, Chief Revenue Officer, Graebel

The Service Order Agent is live today and designed to scale across more than 30 relocation service categories. So far, results include a meaningful reduction in manual effort, faster service-order turnaround, more consistent data quality, and a repeatable blueprint for bringing intelligent automation to the rest of their operations.

Connect intelligent automation systems with Work IQ and interoperable agents

As automation systems become more adaptive, another challenge quickly emerges: connection. Even the most capable agent can only go so far if it operates in isolation. Many organizations are still navigating fragmented ecosystems where agents, workflows, APIs, and external tools all function separately. This naturally makes it difficult to share context or complete work across systems without custom integration effort.

This fragmentation—which is very common—slows adoption and makes it harder to scale intelligent automation beyond isolated pilots.

Now, new interoperability and extensibility capabilities in Work IQ help organizations build more connected agent systems—making it easier for agents, workflows, and enterprise tools to operate together across environments.

With the new Work IQ REST API and command-line interface (CLI) capabilities, teams can integrate Work IQ more flexibly into existing operational and development workflows. Support for remote model context protocol (MCP) servers also introduces a more standardized way to connect agents with tools, services, and enterprise resources. This reduces the need for one-off integrations across growing agent ecosystems.

And as organizations deploy more specialized agents across departments and business processes, coordination between those agents becomes increasingly important. With agent-to-agent (A2A) communication now generally available in Copilot Studio, agents can exchange information, delegate tasks, and be set up to work together more effectively across systems and workflows.

Together, these updates continue the shift from isolated AI experiences toward connected operational platforms—where workflows, agents, APIs, and enterprise systems can be designed to collaborate more naturally across the organization.

Learn more about these new Work IQ interoperability and extensibility updates.

Bring more natural, responsive experiences to customer voice interactions

For many organizations, voice support is still one of the most difficult channels to modernize. Customers get stuck in rigid phone trees, repeat information multiple times, or lose context entirely when they’re transferred to a live agent. Meanwhile, service teams are under pressure to handle growing call volumes without sacrificing customer experience.

Copilot Studio is helping organizations move beyond those limitations with real-time voice agents, now generally available in North America through Dynamics 365 Contact Center. These capabilities help organizations build more natural voice experiences that can identify callers, answer questions, take action during conversations, and transition customers to live agents while preserving context.

Support for speech-to-speech (S2S) voice experiences also makes it easier to connect voice agents into existing customer service and operational systems.

Now, as voice agents become part of real customer interactions, governance and operational readiness become increasingly important. A poor escalation experience, missing context during handoff, or lack of monitoring can quickly become a customer trust issue—not just a technical issue.

That’s why we’re also sharing a new in-depth voice agent governance guide. It covers practical considerations for scaling customer-facing and real-time voice agents responsibly, including escalation testing, monitoring, security, compliance, and operational readiness.

Learn more about real-time voice agents in Copilot Studio and explore the new governance guidance for customer-facing voice agents.

What else is new and improved in Copilot Studio

  • A new orchestration layer in Copilot Studio improves how agents execute business processes with greater accuracy and efficiency. Built on an upgraded AI stack, it has demonstrated measurable gains—improving evaluation performance by approximately 20% while decreasing net token usage by 50%—so agents can complete tasks more reliably and cost-effectively.1 By strengthening tool orchestration and execution quality, the new orchestrator helps ensure more consistent outcomes across complex, multi-step business processes. The result is faster, more dependable automation that scales across enterprise scenarios. This feature is currently in early release environments, where it applies automatically.
  • Agent lifecycle visibility updates help the whole team understand an agent’s approval and publishing status. Agent creators can more easily understand whether an agent is still generating, ready for testing, successfully published, or encountering issues, which reduces guesswork during development and iteration. For IT and platform teams managing agents across environments, clearer publishing and status visibility can make it easier to identify stalled deployments, troubleshoot operational issues, and maintain better oversight as agent programs scale.

All together, these updates are focused on helping teams solve the kinds of problems that slow automation efforts down in the real world: workflows that break when systems change, disconnected tools that create extra manual work, and customer experiences that still feel rigid or fragmented.

With new investments across computer-using agents, workflows, interoperability, and real-time voice, Copilot Studio continues to expand as the agentic platform for building agents, apps, and workflows. The goal is simple: help organizations build systems that are easier to connect, easier to adapt, and easier to operate at scale—without losing the structure, visibility, and governance enterprise teams depend on.

Stay up to date on all things Copilot Studio

More is coming across voice channels, workflows, and the building experience. Check out all the updates as we ship them, as well as new features releasing in the next few months here: What’s new in Microsoft Copilot Studio.

To learn more about Microsoft Copilot Studio and how it can transform productivity within your organization, visit the Copilot Studio website or sign up for our free trial today.

Take your automation further

Design adaptive AI systems that integrate across tools and workflows to help teams get meaningful work done.

Two business professionals looking at the screen of a tablet and collaborating.

1 Source: Microsoft usage data, 2026.

The post New and improved: Computer-using agents, a new workflows experience, and real-time voice experiences appeared first on Microsoft Copilot Blog.

]]>
The in-depth guide to managing real-time voice agents at scale http://approjects.co.za/?big=en-us/microsoft-copilot/blog/copilot-studio/the-in-depth-guide-to-managing-real-time-voice-agents-at-scale/ Tue, 19 May 2026 16:00:00 +0000 Explore an in-depth guide to managing customer-facing real-time voice agents with Copilot Studio, from governance foundations to production readiness.

The post The in-depth guide to managing real-time voice agents at scale appeared first on Microsoft Copilot Blog.

]]>

Governance built into the foundation of your agent program is what separates a successful production deployment from one that stalls—or fails publicly. This guide explains how to design, manage, and scale customer-facing, real-time voice agents using Microsoft Copilot Studio, with a focus on governance, reliability, and enterprise readiness.


Imagine a customer calling your contact center about a billing dispute. A real-time voice agent answers, identifies the customer, references their account history, resolves the issue, and—when needed—hands off to a live agent with full context preserved. Human agents focus on exceptions, not routine queries.

Now imagine that same scenario without agent governance. The agent was built, published directly to production, and never tested for escalation. Monitoring was not enabled. The first signal of a problem is a customer complaint—or a data exposure.

Customer-facing agents are becoming the front door for how organizations engage with customers, handling intent and outcomes across conversational AI experiences. What began as chat has evolved into always-on agents that resolve issues, take action, and now support real-time voice across digital and contact center environments using platforms like Copilot Studio. The opportunity is massive—but so is the cost of getting the foundation wrong. Just as self-service and Q&A agents redefined support at scale, this shift will fundamentally reshape how companies operate.

Why real-time voice agents require a different governance lens

Most organizations already govern internal AI tools designed for known users and controlled environments. Customer-facing agents operate under fundamentally different conditions. There are unknown users, public channels, brand exposure, and direct access to customer data and downstream systems. Failures in these customer experience events mean operational, regulatory, and reputational consequences.

This is why governance cannot be treated as a final approval step. As real-time voice agents scale, governance must be built into how they are designed, deployed, monitored, and evolved from the start. Organizations that treat governance as an accelerant—rather than a constraint—can move faster and more confidently than those who bolt it on later.

Principle: Governance as a design principle can streamline approval, which leads to accelerated scale and adoption.

Why real-time voice agents raise the stakes

Text‑based agents require governance, but real‑time voice introduces stricter operational constraints. Latency budgets are tighter, failures are immediately apparent to customers, and interruption handling, turn‑taking, session state, and escalation behavior directly affect service reliability.

Voice agents are typically deployed in high‑impact scenarios such as billing, orders, and service disruptions, where they integrate with Dynamics 365 Contact Center workflows. In these environments, agents must identify callers, reference active cases, execute actions, and escalate predictably.

For real‑time voice, escalation is a first‑class system requirement. Handoffs to human agents must preserve full conversational context and session state, and be validated under load before production traffic is routed.

Model selection also becomes operationally significant. Copilot Studio real‑time voice agents use purpose‑fit models to balance latency, quality, and reliability while remaining governed through a centralized control plane.

What good looks like: A production voice agent deployment has been tested for escalation behavior, latency under load, and handoff context preservation before any customer traffic is routed to it. Monitoring is active from day one, not added after the first incident.

A governance framework for the full agent lifecycle

Governing customer-facing agents effectively requires capabilities that span the full agent lifecycle. This is especially critical for business-to-consumer (B2C) agents, which operate in always-on, customer-facing contexts and must handle real-time interactions, actions, and sensitive data at scale—particularly in high‑stakes modalities like voice.

Copilot Studio provides this governance as a managed agent platform, enforcing controls through managed operations and managed security across the full lifecycle. That goes from build access and data connectivity to release, monitoring, and auditability. Rather than relying on documentation or custom wiring, governance is centralized in the Microsoft Power Platform control plane and consistently applied across chat, voice, and contact center scenarios.

The following five‑stage governance framework reflects how managed capabilities come together across the full lifecycle of customer-facing agents:

  1. Govern the builder
  2. Govern the build
  3. Govern the release
  4. Govern the runtime
  5. Govern the lifecycle

Stage 1: Govern the builder

Before a single topic is created, agent governance starts with who is allowed to build and what they are allowed to connect.

  • Define builder roles and environments. Specify who can create agents and which environments they can work in, using role‑based access in the Power Platform admin center.
  • Set data access boundaries early. Apply data loss prevention (DLP) policies before development to determine which connectors and data sources agents can use.
  • Maintain environment separation. Use distinct development, test, and production environments to validate changes before deploying them to customer‑facing scenarios.
  • Standardize on managed solutions. Package agents in managed solutions to support versioning, controlled promotion, and rollback across environments.

What good looks like: A new agent builder requests access and is provisioned into a dedicated development environment. DLP policies are pre-applied. They cannot publish to any customer-facing channel without an administrator approval step.

Stage 2: Govern the build

How an agent is built determines how safe and predictable it is in production.

  • Configure authentication by channel. Decide whether sessions are authenticated (Microsoft Entra ID or supported identity provider [IdP]) or anonymous, and design data access accordingly. (For public-facing scenarios like 800 numbers and public websites, anonymous real-time voice sessions are common.)
  • Set generative AI behavior explicitly. Define and check grounding, topic scope, and allowed behaviors rather than relying on default settings.
  • Validate escalation paths. Test and verify handoff to live agents with full conversation context preserved for all voice scenarios.
  • Apply content moderation intentionally. Define clear engagement boundaries, enforce agent governance and policy controls, and rigorously red‑team and validate edge cases before deploying to production.

What good looks like: Testing escalation paths before publishing an agent to a customer-facing channel, so you can go live with more confidence. Catching errors before the first live escalation is critical to creating a good customer experience.

Stage 3: Govern the release

Moving an agent from development to production requires controlled, auditable steps.

  • Standardize promotion paths. Promote agents through dev, test, and production using managed solutions and Power Platform pipelines with an auditable change history.
  • Apply preproduction validation gates. Require checks for conversation quality, escalation behavior, latency under load, and data access before publishing.
  • Plan and test rollback. Define and validate rollback procedures for production issues prior to go‑live.
  • Separate publish authorization. Require explicit approval to publish agents to customer‑facing channels, independent of build permissions.

What good looks like: An agent must pass a defined pre-production checklist and receive administrator approval to publish before any customer traffic reaches it. Every version promotion is tracked in the solution history.

Stage 4: Govern the runtime

Once an agent is live, governance shifts from control to visibility and response.

  • Enable runtime observability. Turn on conversation transcripts and analytics in Copilot Studio before routing customer traffic.
  • Define operational thresholds. Monitor metrics such as escalation rate, resolution rate, latency, and session completion, with alerts for deviations.
  • Establish incident response. Define processes for detecting, triaging, and mitigating production issues in voice agents integrated with Dynamics 365 Contact Center.
  • Monitor usage and capacity. Track session volume, message usage, and capacity limits to support scaling and stability.

What good looks like: Early detection through active monitoring. Voice agents that interact with customers without active monitoring are operating without a safety net. Issues that could persist for hours without analytics can be caught in minutes with these guards in place.

Stage 5: Govern the lifecycle

Voice agents are not static. They evolve as scenarios expand, customer needs change, and the platform advances. Managing change safely is as important as the initial deployment.

  • Version agent configuration. Track changes to topics, actions, authentication, and generative AI settings using application lifecycle management (ALM) and source control.
  • Validate changes preproduction. Test all updates in non‑production environments to avoid regressions in core scenarios, including voice flows and escalation behavior.
  • Coordinate releases operationally. Communicate deployment windows to IT and contact center operations teams.
  • Evolve governance as scale grows. Reassess role-based access control (RBAC), DLP policies, environment strategy, and publishing permissions as agent count and channel coverage expand.

Platform capabilities that support agent governance

Copilot Studio provides a centralized control plane for building, operating, and governing customer‑facing agents. The platform capabilities below directly enable the governance framework described above and should be configured before scaling B2C deployments:

  • Power Platform admin center: Central governance surface for environments, DLP policies, user access, and capacity management; the primary enforcement layer for agent governance.
  • Environment management: Separate development, test, and production environments to support validation and controlled promotion of customer‑facing agents.
  • Data loss prevention (DLP) policies: Environment‑level connector controls that define which data sources and services agents can access before any connections are established.
  • Managed solutions and Power Platform pipelines: Package agents as managed solutions and promote them through environments with version tracking, rollback support, and an auditable change history.
  • Microsoft Entra ID and channel authentication: Configure customer‑facing authentication using Entra ID or supported identity providers to enable secure, scoped access to customer data.
  • Generative AI controls and content moderation: Per‑agent configuration for grounding, topic scope, allowed behaviors, and content filtering, applied deliberately prior to public deployment.
  • Conversation transcripts and analytics: Built‑in logging and analytics providing runtime visibility into agent behavior, escalation patterns, and coverage gaps.
  • Dynamics 365 Contact Center integration: Native escalation to live agents with case context preservation and unified conversation history for voice deployments.
  • Azure Speech: Underlying speech infrastructure for real‑time voice agents, with implications for latency, reliability, and capacity planning.
  • Dataverse security model: Row‑level and business‑unit security controls governing agent access to customer records in Dynamics‑integrated scenarios.

Security, privacy, and compliance for customer-facing agents

For IT and security teams, governance of customer-facing agents must also address data handling, regulatory requirements, and audit readiness. These are not secondary concerns—they’re often the first gate any enterprise B2C deployment must pass through.

Customer data and PII in voice interactions

Real-time voice agents generate conversation transcripts that may contain personally identifiable information. Establish clear retention policies for these transcripts before deployment. Define who has access to conversation logs, how long they are retained, and whether they are subject to deletion requests under applicable privacy regulations.

Regulatory considerations

Depending on your industry and geography, customer-facing AI agents may be subject to requirements under General Data Protection Regulation (GDPR), California Consumer Privacy Act (CCPA), or sector-specific regulations in financial services or healthcare. Review applicable requirements with your legal and compliance teams before deploying agents to regulated customer scenarios. DLP policies in the Power Platform admin center are a key compliance control.

Audit logging and compliance evidence

Power Platform and Copilot Studio support audit logging through Microsoft Purview and the Power Platform admin center. Ensure audit logging is enabled before production deployment and that logs are retained according to your organization’s compliance requirements.

Credential and secret management

Agents that connect to external systems require credentials and connection strings. Do not store secrets in agent configuration directly. Use environment variables in Power Platform or Azure Key Vault references to manage credentials securely, with access controlled through role assignments.

Note for architects: Security and compliance review should be a gate in Stage 3 (govern the release), not an afterthought discovered during audit. Engage your security and compliance teams in the pre-production validation checklist.


Five anti-patterns that derail production AI deployments

Organizations that have scaled B2C agents successfully tend to have avoided the same set of avoidable mistakes. These are the patterns most likely to cause problems once customer traffic is live.

  1. Skipping environment separation: Building and publishing agents in the same environment, or directly in production, allows untested changes to reach customers and is one of the most common causes of early deployment issues.
  2. Publishing voice agents without tested escalation: Escalation to a live agent is a core part of voice agent design. Untested handoff paths that fail to preserve customer context degrade the experience more than having no agent at all.
  3. Granting broad DLP exceptions under schedule pressure: Temporarily relaxing DLP policies often becomes permanent, introducing data access risk and audit gaps that are difficult to remediate later.
  4. Treating monitoring as a postlaunch activity: When transcripts, analytics, and alerts are not enabled before go‑live, production issues surface through customer complaints rather than operational signals.
  5. Building openended agents without defined scope: Broad, general‑purpose agents are harder to test, govern, and improve than agents scoped to specific customer scenarios with clear success criteria.

How to operationalize voice agents

As teams move from pilots to production, a small set of patterns consistently differentiates voice agent deployments that scale.

  • Start with well‑defined customer scenarios rather than broad open‑ended agents. Clear scope simplifies risk assessment, testing, and measurement. A voice agent designed for order status or billing inquiries is easier to govern and iterate on than one intended to answer arbitrary customer questions.
  • Treat real‑time voice as an extension of existing digital agent governance, not an exception. Teams that have already governed chat‑based agents in Copilot Studio are well positioned to apply the same controls to voice, while accounting for stricter latency, escalation, and runtime requirements.
  • Design escalation as a primary flow, not a fallback. Agents integrated with Dynamics 365 Contact Center should preserve full conversational and case context on handoff. Predictable escalation maintains continuity; dropped context undermines trust.
  • As programs scale, three governance questions remain central:
    • Which customer scenarios are appropriate for automation versus human handling?
    • Where does real‑time voice materially improve the experience versus add operational complexity?
    • How quickly can production issues be detected and resolved once agents are live?

Using Copilot Studio as a governance foundation for agents

Copilot Studio and Power Platform provide a centralized environment for building, operating, and governing agents, which becomes increasingly important as deployments expand from internal use cases to customer‑facing channels.

Establish governance once in Copilot Studio, and scale it across chat, voice, and backend‑driven agents without fragmentation. As a centralized control plane, the platform helps you enforce consistent policies and maintain operational oversight as agents expand across channels, regions, and customer scenarios.

For organizations already using Copilot Studio, many of the governance capabilities described here are available today. Support for real-time voice agents in Copilot Studio is now generally available in North America, with deployments delivered first through Dynamics 365 Contact Center. Language support, additional regions, and broader publishing channels will expand over time as part of Copilot Studio’s ongoing roadmap.

Learn more in the announcement blog for real-time voice agents.

Governance readiness checklist for customer-facing voice agents

Before deploying a customer-facing or real-time voice agent to production, verify governance readiness across these core dimensions.

Access and environment

  • Separate development, test, and production environments are provisioned
  • Role-based access is configured—developers cannot publish directly to production
  • Advanced connector policy is applied to all environments before development begins
  • Publishing permissions for customer-facing channels require administrator approval

Build and configuration

  • Authentication and identity are configured appropriately for the channel (authenticated or anonymous)
  • Generative AI settings, grounding, and content moderation are configured deliberately
  • Credential and secret management uses environment variables or Azure Key Vault references
  • The agent is packaged in a managed solution with tracked versioning

Testing and release

  • Escalation paths to live agents have been tested with context preservation verified
  • Latency and behavior have been validated under simulated load
  • A pre-production validation checklist has been completed and signed off
  • A rollback procedure has been defined and tested
  • Audit logging is enabled and log retention meets compliance requirements

Runtime and operations

  • Conversation transcripts and analytics are active before first customer interaction
  • Operational thresholds (escalation rate, session completion rate) are defined with alerts
  • An incident response procedure is defined and communicated to operations teams
  • Usage monitoring is in place for capacity planning
  • A change management process is defined for updating live agents

Getting started with customer-facing agents

Organizations ready to operationalize B2C agents should begin with the following steps:

  • Align on priority scenarios. Agree on customer scenarios, scope, success criteria, and escalation requirements before any development begins.
  • Set up environments and governance. Configure separate dev, test, and production environments and apply DLP policies before granting developer access. Define role‑based access and require administrator approval for publishing to customer‑facing channels.
  • Engage security and compliance early. Review applicable regulatory requirements and establish data retention policies for conversation transcripts.
  • Build and validate deliberately. Start with a scoped agent, use managed solutions, and be sure to test and verify escalation paths.
  • Confirm readiness before golive. Complete the governance readiness checklist and enable monitoring and escalation thresholds prior to routing customer traffic.

With the right foundation in place, teams can scale customer‑facing and real‑time voice agents—while maintaining the reliability, security, and operational integrity IT teams are responsible for protecting.

Resources for governing AI agents

Governance starts early

Establish a governance foundation with Copilot Studio that scales across chat, voice, and backend-driven agents.

A person working on a computer in an open office setting.

The post The in-depth guide to managing real-time voice agents at scale appeared first on Microsoft Copilot Blog.

]]>
New and improved: Agent governance, intelligent workflows, and connected app experiences http://approjects.co.za/?big=en-us/microsoft-copilot/blog/copilot-studio/new-and-improved-agent-governance-intelligent-workflows-and-connected-app-experiences/ Mon, 11 May 2026 16:00:00 +0000 See what's new in Copilot Studio, April 2026: updates to workflows, increased control over agent operations, and an expanded agent usage estimator.

The post New and improved: Agent governance, intelligent workflows, and connected app experiences appeared first on Microsoft Copilot Blog.

]]>

As organizations scale their use of AI agents, IT teams face a familiar challenge: how do you expand automation without losing control? Individual agents can be powerful, but as they connect through workflows and integrate across systems, requirements for visibility, governance, and predictability become much more complex. And capability must be grounded in confidence.

The April 2026 updates in Microsoft Copilot Studio focus on building that confidence across the platform. From increasing visibility and governance for admins to expanding intelligent workflow capabilities, these features help you move from isolated automation to connected, reliable systems.

Build and scale agents with better visibility and control

As agents expand across organizations and business processes, admins need clear visibility into how they’re performing, how they’re secured, and what they’ll cost to run. These updates help you manage agents more effectively without adding more friction—or risk.

See agent performance and status more clearly

Copilot Studio now surfaces agent status directly in the authoring experience, giving you immediate insight into each agent’s security and protection posture. You can quickly identify issues like authentication gaps or policy impacts and investigate them at the source. This helps reduce guesswork and speed up resolution.

As you gain clearer visibility into agent performance, you can also share those insights more safely. The Analytics Viewer role, now generally available, introduces read-only access to an agent’s Analytics page.

The Analytics Viewer role allows us to provide meaningful performance insights to business and operational stakeholders while maintaining strict production governance. It cleanly separates operational visibility from agent configuration and publishing rights.

—Mohamed Arhab, Solution Architect, City of Montreal

Allowing analysts and stakeholders to monitor performance, without giving them the ability to modify the agent, helps resolve a long-standing tradeoff between visibility and control. Now it’s easier to share insights broadly while maintaining clear separation of responsibilities.

Speaking of extending visibility and control, there’s more good news: Microsoft Agent 365 is now generally available. Agent 365 is the centralized control plane for managing agents across your environment. This brings together visibility into agent inventory, permissions, behavior, and activity in one place so that you can monitor and govern agents consistently, not just where they’re built.

For Copilot Studio customers, this means the agents you create can be managed alongside agents from Microsoft 365 and partner ecosystems, with shared policies, security controls, and lifecycle oversight. As Agent 365 continues to expand its integrations and multi-agent capabilities, it further strengthens Copilot Studio’s role as the place where agents are built—while governance scales across the full system. Learn more about Agent 365.

Plan and scale with clearer cost visibility

The expanded agent usage estimator now includes Dynamics 365 agents, such as Sales Qualification Agent and Customer Service Agent. By forecasting Copilot credit usage across both Copilot Studio and Dynamics 365 scenarios in one place, you can model usage more accurately and scale deployments—helping avoid unexpected cost surprises.

With these recent admin updates, the result is fewer bottlenecks, better-informed decisions, and a clearer path to scaling agents across your organization.

Expand workflows into intelligent, governed automation systems

In Copilot Studio, workflows are step-by-step automation processes that complete actions or tasks in a deterministic, reliable way. As workflows become the backbone of business automation, these new updates help you extend their capabilities—bringing in more AI-powered reasoning, centralized governance, and a growing ecosystem of tools in a way that’s reliable and secure by design.

Design and validate workflows with more clarity

One powerful way to make your workflows more adaptable and effective is by embedding Copilot Studio agents directly into them. Using agent nodes inside workflows means that instead of just performing the task with rigid logic, the workflow can delegate reasoning, decisions, or output generation to an agent at any prescribed step of the process.

This makes workflows more resilient to real-world situations—which have a lot of variability—while still following the defined structure that make IT teams less nervous.

In addition to embedding agents, you can now also add and configure AI actions directly within the flow to understand requests, route work, and generate content dynamically. And with the ability to test individual steps using sample inputs, teams can validate behavior earlier, debug more effectively, and refine workflows before they’re deployed.

In practice: Unifi, North America’s largest provider of aviation ground handling services, used Copilot Studio and Power Platform to automate legal contract review by combining agents with deterministic workflows. Instead of relying on a single agent, they broke the process into coordinated steps that extract, classify, and validate key terms across documents. This system reduced contract processing from days to minutes and delivers the same level of performance as much more expensive, off-the-shelf products built specifically for the legal industry.

The result is a workflow experience that’s more adaptable and more predictable to operate. This helps give teams—both makers and administrators—more confidence in creating more sophisticated automation that doesn’t sacrifice clarity or control.

Scale workflows across systems with built-in governance

Speaking of clarity and control, there are also new updates to workflows that help you scale automation without introducing new governance risks.

Workflows can now connect to a broader ecosystem of tools, including model context protocol (MCP) server-enabled tools (preview), which makes it easier to take action across systems while staying within Microsoft security, permission, and compliance boundaries. This allows workflows to execute tasks and involve users for review and approval within governed processes.

We’ve also introduced a centralized, admin-controlled environment for Workflows Agent. This makes it easier to apply data loss prevention (DLP) policies consistently and maintain visibility across automation, so workflows remain compliant by design, even as they scale.

Together, these updates make it easier to move from isolated automations to connected, intelligent systems. With those systems, you can scale workflows across your organization with greater confidence, control, and flexibility.

Bring business apps directly into your agents

As agents become part of everyday work, a common gap emerges: they can generate insight, but acting on that insight often requires switching tools, re-creating context, or handing work off across systems. Support for apps in agents, now generally available, helps to close that gap.

Turn intent into action inside Copilot Chat

Agents built in Copilot Studio can now surface rich, interactive app experiences directly in Copilot Chat, allowing users to review data, update records, approve requests, or create assets in place. Instead of switching tools or re-creating context, work happens seamlessly within the flow of conversation. This helps reduce friction and empowers teams to move faster from insight to execution.

Animated UI showing Adobe Express embedded in Microsoft 365 Copilot chat, where a user accesses design templates and visuals directly within the conversation.

Work across the systems your business already runs on

Apps in agents bring together Microsoft and partner applications—from Power Apps to Dynamics 365 and beyond—so agents can take action across the systems your teams already use. These experiences are built and orchestrated in Copilot Studio, where you define how agents interact with apps, data, and workflows to support real business processes.

Extend and scale with trusted integrations

Through the Agent Store, you can adopt ready-made agent experiences or extend your own with partner-built integrations—while maintaining enterprise-grade security, permissions, and admin control. Options include:

  • Adobe Express (seen above)
  • Box
  • Figma
  • Monday.com
  • Wix

These options (and more) make it easier to scale agent usage across your organization without losing oversight.

These capabilities, all generally available now, help teams shift agents from being informational tools to operational ones. They bring real business actions into Copilot Studio agents in a way that’s both more functional for users and manageable for IT—helping teams complete work efficiently while maintaining the governance needed to scale.

Learn more about apps in agents.

What else is new and improved in Copilot Studio

  • Evaluation insights and automation updates now make it easier to generate test cases from analytics, simulate multi-turn interactions, and automate evaluations through APIs and connectors. You can turn real user conversations into targeted test sets, better reflect complex, real-world scenarios, and run evaluations programmatically. Together, these capabilities help you operationalize agent quality and maintain confidence as you scale.
  • Custom metrics for outcome-based measurement help you track what actually matters to your business, not just usage. Define success in your own terms—like resolution rates or conversions—and automatically evaluate conversations against those outcomes, making it easier to understand impact, align stakeholders, and make data-driven decisions.
  • Work IQ API is now available in public preview to bring Copilot’s intelligence layer—grounded in organizational context, memory, and signals—into your own agents and workflows. With built-in orchestration and enterprise-grade security, you can build agents that understand what’s happening across your business without managing raw data or complex integrations.
  • Agent-to-agent (A2A) communication is now supported in Work IQ, allowing agents to collaborate as peers and delegate tasks using shared organizational context. This makes it easier to build multi-agent systems that can coordinate work, maintain context across interactions, and deliver more grounded, role-aware outcomes.
  • GPT-5.5 Thinking is now available in Copilot Studio early release cycle environments as GPT-5.5 Reasoning, further expanding model choice with its more advanced analysis capabilities. This model is also rolling out across Microsoft 365 Copilot in Copilot Chat, Word, Excel, and PowerPoint.

Stay up to date on all things Copilot Studio

More is coming across voice channels, workflows, and the building experience. Check out all the updates as we ship them, as well as new features releasing in the next few months here: What’s new in Microsoft Copilot Studio.

To learn more about Microsoft Copilot Studio and how it can transform productivity within your organization, visit the Copilot Studio website or sign up for our free trial today.

Build agents your way

Create, deploy, and scale custom agents and workflows with Copilot Studio.

A person working on a laptop.

The post New and improved: Agent governance, intelligent workflows, and connected app experiences appeared first on Microsoft Copilot Blog.

]]>
Extend AI voice support: Introducing real-time voice agents in Microsoft Copilot Studio http://approjects.co.za/?big=en-us/microsoft-copilot/blog/copilot-studio/extend-ai-voice-support-introducing-real-time-voice-agents-in-microsoft-copilot-studio/ Mon, 27 Apr 2026 15:00:00 +0000 Real-time voice agents are now generally available in Copilot Studio, supporting adaptive voice experiences for complex customer conversations.

The post Extend AI voice support: Introducing real-time voice agents in Microsoft Copilot Studio appeared first on Microsoft Copilot Blog.

]]>
Customers expect support that resolves issues quickly, delivers consistent answers, and works seamlessly across channels. For organizations, this creates a familiar tension: how do you deliver high‑quality service at scale without losing control over cost, compliance, or experience?

That’s why we’re excited to announce the general availability of real‑time voice agents in Microsoft Copilot Studio launching in Dynamics 365 Contact Center. These agents are designed for nuanced, high‑impact interactions where voice experiences need to adapt in the moment—while still operating within trusted, enterprise‑grade solutions.

Real-time voice agents build on the momentum of Copilot Studio as a proven, enterprise-scale platform. Over 80% of Fortune 500 companies now have active agents built using our low-code/no-code tools, creating a strong foundation for bringing real-time, conversational voice experiences into their customer workflows.1 This widespread adoption shows how ready organizations are to extend their existing Copilot-powered agents into natural, responsive voice interactions.

Why AI voice support needs to go off script

For decades, menu‑based interactive voice response (IVR) systems provided predictability, reliability, and compliance at scale. Over time, organizations layered in speech recognition and automation to reduce friction and manage growing call volumes more efficiently. These approaches remain critical today and continue to power successful customer service operations across industries.

What’s changed is not the importance of voice, but the expectations customers now bring into voice interactions. Conversations rarely follow a straight line. Customers interrupt, clarify, change direction mid‑call, or introduce urgency without warning. When that happens, rigid interaction models can struggle to keep up, even when the underlying systems are reliable.

Meanwhile, across industries, contact centers are under pressure to do more with less. Interaction volumes continue to rise, margins are tighter, and customer expectations are shaped by digital experiences that feel fast, personal, and responsive. Many teams are already exploring AI through chatbots, voice automation, or workflow tools, but scaling those experiments into production voice experiences that customers trust is far harder than running a pilot.

Voice, in particular, exposes gaps immediately. Latency, awkward handoffs, or missing context are noticed in real time, often at the most critical moments. That’s why simply adding automation isn’t enough. As organizations move toward agentic, AI‑first service models, voice needs to work as part of a unified service layer: one where understanding, reasoning, and action happen together, and where context carries forward if escalation is required.

Meeting this bar doesn’t require replacing the voice systems that already work. It requires extending them with capabilities designed specifically for live, conversational interactions. This makes it possible to move beyond scripted interactions without treating voice as an isolated AI experiment.

Meet real-time voice agents for AI voice support

Real‑time voice agents represent a new premium mode within voice agents that sits under the broader category of conversational AI. They are optimized for low‑latency, interruptible, speech‑to‑speech conversations with real‑time reasoning. This distinction helps teams choose the right interaction model for each scenario without treating all voice interactions as equal.

These agents can move from intent to action to confirmation within a single interaction. They can retrieve or update information mid‑conversation and take action as the interaction unfolds. Customers can get faster resolution, while service teams can maintain consistent experiences and centralized operational control. This also frees human agents to focus their time and energy on conversations where judgment and empathy matter most.

High-volume self-service, designed for everyday voice interactions

Real-time voice agents are built for the high volume, business-to-consumer (B2C) interactions that define customer service at scale. These everyday inbound calls make up the majority of customer engagement across industries. To support this breadth, Copilot Studio enables organizations to grow their voice strategy—starting with deterministic, template-driven flows and expanding into dynamic, real-time voice agents as needs evolve.

Copilot Studio provides a documented set of external voice agent templates designed for unauthenticated, customer-facing interactions across both Copilot Studio and Dynamics 365 Contact Center, covering many of the core workflows customer service teams handle today, including:

  • Billing and payments, which are among the most frequent and sensitive contact center interactions. Customers want clarity and resolution in the same call, not a handoff or follow-up. These scenarios often start with deterministic flows—confirming identity, checking balances, or processing payments—but can switch to dynamic, real-time voice agents that explain charges, respond to questions, and adapt tone as situations become more urgent.
  • Order and reservation support across retail, travel, and hospitality, which often requires more dynamic handling. What begins as a simple status check can quickly shift into a change request or issue resolution. Real-time voice agents adapt to these pivots by grounding responses in live order or reservation data. They’re designed to act while preserving relevant context, supporting organizations as they evolve from structured flows to more flexible conversational automation.
  • Eligibility and verification scenarios in healthcare, financial services, telecom, and the public sector, which rely on accuracy and trust. Deterministic steps—collecting required information, confirming eligibility criteria—can be paired with dynamic voice capabilities that answer questions as they arise and help preserve continuity if escalation or additional support is needed.
  • Appointment scheduling and changes often involve interruptions and evolving preferences. Organizations can begin with predictable scheduling flows and grow into dynamic, real-time voice agents that let customers book, reschedule, or confirm details conversationally, without unnecessary restarting or repeating information as needs shift.
  • Account and membership management, which often spans multiple tasks, from updating personal details to managing subscriptions or reviewing benefits. Deterministic updates can be combined with dynamic, context-aware voice interactions that stay connected to live account data, keeping these conversations efficient, accurate, and natural from start to finish.

From conversation to resolution, without breaking flow

Real‑time voice agents are built for real customer service environments, where escalation is a natural part of resolution rather than a failure of automation. At launch, these experiences are delivered through Dynamics 365 Contact Center, with conversation context carrying forward automatically, helping reduce the need for customers to restate information when human judgement is required.

As part of the Copilot Studio product roadmap, real‑time voice agents will also extend to Microsoft Teams Phone and additional Copilot Studio digital apps and channels, enabling organizations to bring consistent, context‑aware voice experiences to more customer touchpoints over time.

Voice automation designed for trust, control, and evolution at scale

For IT teams, voice automation has always been high impact but high risk. Feedback from early internal deployments reinforces the importance of pairing conversational intelligence with enterprise‑grade governance and lifecycle controls.

Innovation in voice does not require giving up control or predictability. Voice is a high‑stakes interaction surface, and customers notice immediately when experiences feel unreliable or inconsistent.

Copilot Studio takes a deliberate approach to real‑time voice by selecting the right models for the right moments, including the latest frontier models, optimized for the right balance of quality, latency, and reliability, without placing that complexity on customers. Built on Microsoft’s enterprise foundations for security, governance, and operational oversight, real‑time voice agents help organizations innovate with confidence while preserving trust.

Get started with Copilot Studio

Support for real-time voice agents is now generally available in North America for Dynamics 365 Contact Center, with language support, additional regions, and broader customer touchpoints expanding over time as part of Copilot Studio’s global rollout.

Learn more about real-time voice agents and sign back in to Copilot Studio to start scaling customer support with agents today.

New to Copilot Studio? Discover how you can transform your business by building, evaluating, managing, and scaling custom AI agents—all in one place. Sign up for a free trial of Copilot Studio today.

Build voice agents that adapt

Create and manage high-impact, flexible voice agents—designed for enterprise scenarios.

IT professionals gather around a computer in a modern open office space.

1 Source: Microsoft usage data, 2026.

The post Extend AI voice support: Introducing real-time voice agents in Microsoft Copilot Studio appeared first on Microsoft Copilot Blog.

]]>
Automate business processes with agents plus workflows in Microsoft Copilot Studio http://approjects.co.za/?big=en-us/microsoft-copilot/blog/copilot-studio/automate-business-processes-with-agents-plus-workflows-in-microsoft-copilot-studio/ Fri, 10 Apr 2026 15:58:12 +0000 Introducing new capabilities in Microsoft Copilot Studio that help you automate your business processes by mixing AI agents and workflows.

The post Automate business processes with agents plus workflows in Microsoft Copilot Studio appeared first on Microsoft Copilot Blog.

]]>
Today we are introducing new capabilities in Microsoft Copilot Studio that help you automate your business processes by mixing AI agents and workflows. Agents and workflows already exist in Copilot Studio as two complementary capabilities with unique strengths. Agents bring reasoning and adaptability; workflows bring structure and consistency.

So how do you know when to use agents vs. workflows?

It’s no longer an either-or decision. Here’s how to use agents and workflows together to combine strengths and reduce risks.

What are agents and workflows?

Agents are flexible AI solutions that rely on foundational models to act, share knowledge, and handle tasks. They are powerful precisely because they are flexible. They can interpret unstructured inputs, reason over context, and make decisions beyond fixed logic.

However, organizations often need to know that repetitive parts of their processes will behave consistently, every time they run. Pure agent autonomy doesn’t always hold up to that requirement in production.

Screenshot of the Copilot Studio homepage, showing options to create a workflow or create an agent

Workflows, by contrast, are powerful automations that drive process execution with consistency and speed. They’re designed to deliver the reliability that many business processes require.

At the same time, rigid, rules-based automation has its own ceiling. It’s nearly impossible to anticipate every potential input format, edge case, and decision-making context when building a workflow ruleset. Thus, when the workflow automation encounters something unexpected, it can’t move forward.

Two patterns for scaling automation with AI

While both agents and workflows have their strengths, we’re seeing customers get the most value in Copilot Studio by combining the two. In practice, we’re observing two patterns emerge in how customers apply Copilot Studio, and we’re continuing to deliver product improvements to strengthen and support them.

Workflows that use agents

The first pattern is workflows that call agents. In these instances, the workflow provides the structure for the business process—the defined steps, branching logic, handoffs, and an audit trail. Meanwhile, the agent handles the parts of the process that require judgement. This might include interpreting a document, synthesizing information from multiple sources, or deciding how to route an exception.

Once the agent completes its work, control returns to the workflow, and execution continues predictably.

To make it easier to add agents to workflows in Copilot Studio, we’re introducing agent nodes: the ability for workflows in Copilot Studio to call an agent directly within a workflow. You can build a deterministic, reliable automation, and at the exact moment you need AI reasoning, the flow simply hands it off to an agent.

Setting up an agent node inside a workflow is simple:

  1. Create a workflow step called “Add an agent.”
  2. Select any Copilot Studio agent you’d like to be include in the workflow.
  3. Provide the instructions or task the agent needs to fulfill, and include an option to contact a designated person if specific clarification is needed.
  4. Add the rest of the workflow’s steps.

When you run the workflow, the agent will do its job at the appropriate stage, and then the rest of the workflow will automatically continue.

Screenshot of the Workflow editor, showing a "Run an agent" step and the instructions for calling the agent inside the workflow
Adding an agent node inside a workflow

When to use agents inside workflows

Using agent nodes to include agents in your workflows unlocks scenarios that rigid automation alone can’t handle. Some potential uses include the following:

  • A procurement workflow that routes to an agent to evaluate vendor proposals against company policies.
  • An HR onboarding workflow that personalizes welcome materials based on role and department.
  • A customer service process that escalates complex cases to an AI agent for resolution recommendations.

In general, anywhere your workflow hits a decision that can’t be captured in simple if-then logic—where it needs to use reasoning over context, orchestrate tools, or retrieve knowledge from multiple sources—an agent node can help bridge the gap and make your workflow more effective. This capability is available now in all regions.

Agents that use workflows

The second pattern is equally important: agents that use workflows as tools. When an agent is working through a complex task, it doesn’t need to rediscover how to act every time. Instead, it can call a reliable, tested workflow to execute a well-defined subprocess—and then use the result to continue its reasoning and response.

This ability helps agents to build on existing process infrastructure rather than reinventing it. Moreover, it helps give organizations more confidence that the high-frequency or high-stakes parts of the processes can run with the consistency and controls the org requires.

There are two ways to add workflows into an agent:

  1. Use natural language to build a workflow directly inside Copilot Studio and include that new workflow in an agent.
  2. Alternatively, from within the agent, you can access your library of pre-existing workflows and add them as tools. Then, provide explicit instructions to your agent on when to use the workflow.

That’s it—your agent’s orchestrator will select the right workflows at the right time when needed to complete its work.

Library of pre-existing flows you can add to your agent

When to use workflows inside agents

Adding workflows inside your agents helps add structure and consistency to interactions that still require flexibility. Some potential uses include the following:

  • A sales agent assembles the right product details and pricing tier for a deal, then calls a workflow to generate the quote, apply discount rules, and route it for approval.
  • A customer service agent determines a refund is warranted, then calls a workflow to validate it against business rules, process the payment reversal, and send the confirmation.
  • A procurement agent evaluates which vendor and terms apply to a request, then calls a workflow to create the purchase order in the ERP system and routes it through the approval chain.

Generally, anywhere your agent needs to reliably execute a repeatable process—enforcing business rules, coordinating systems, or ensuring key steps are completed—a workflow can help ground its actions and make outcomes more consistent.

Start using agents and workflows together

Together, these two ways to combine agents and workflows provide you with flexibility to build automations that work better for your real-world needs. Agents handle ambiguity where workflows go brittle; workflows enforce structure where agents might drift.

By embracing a combination of agents and workflows, it becomes easier for different teams to engage in ways that fit the way they work best. Business teams can extend and adapt these automation solutions without rebuilding from scratch. Compliance teams can audit them. Finally, your security and governance teams can choose the right balance of consistency and agility, based on what each scenario requires.

In organizations already using Copilot Studio to support their daily work, both patterns—workflows using agents and agents using workflows—show up regularly:

  • A procurement workflow calls an agent to evaluate supplier contracts that arrive in inconsistent formats.
  • A customer service agent, handling an open-ended request, calls a workflow to initiate a refund or update an account record.
  • An approval process invokes an agent to synthesize context before routing to a decision-maker—and separately, that same agent calls a workflow to send notifications, log outcomes, or kick off downstream steps.

These scenarios show how automation and intelligence can reinforce each other, combining structure and flexibility to deliver more adaptable, dependable results.

Try these capabilities in Microsoft Copilot Studio today.

Transform your business processes

Build powerful AI agents and workflows to assist with and automate work

Two men collaborating joyfully at work

The post Automate business processes with agents plus workflows in Microsoft Copilot Studio appeared first on Microsoft Copilot Blog.

]]>