Sam's Dynamics, Power Platform & AI Blog

Showing posts with label ALM. Show all posts
Showing posts with label ALM. Show all posts

Thursday, 17 September 2026

Testing a Copilot agent: evaluation, not "it seemed to work"



Given input, expect output works for a plugin. Agent responses depend on prompt, retrieval and context interpretation, so you measure response quality and task alignment, not correctness.


Test sets. Built into Copilot Studio. Up to 100 cases per set — hand-written, spreadsheet import, or AI-generated. Quick set: 10 questions from the agent's description and instructions. Full set: up to 100 from knowledge sources or topics.

The generator reaches knowledge sources using the connected account's credentials, so generated cases can contain sensitive data that account can see. Review before sharing the set.

Grading. Every method except general quality needs expected responses or keywords — writing down what a good answer looks like is the actual work. Custom Graders (classification method) encode your own policies where built-in dimensions don't fit.

Test user profiles. Evaluations run under a designated test account, and that account connects to knowledge sources and tools during the run. Testing as yourself proves nothing about what a sales rep sees. Simulate profiles to check behaviour across roles and access levels.

Limits: GCC can't add user profiles to test sets and doesn't support the similarity method. User-auth evaluations need the Copilot Studio connector enabled.

Automation. Results live 89 days — export CSV for anything auditable. REST API triggers evaluations programmatically for release validation and CI/CD regression runs.

Copilot Studio Kit goes deeper: tests via Direct Line API, enriched from Application Insights and Dataverse transcripts, so you get triggered topic and intent recognition scores behind each pass or fail. Multi-turn tests, and pipeline gating — deploy pauses, tests run, thresholds checked, then promote.

What goes in the set:

  • Questions users actually ask, badly phrased ones included
  • Every bug that ever shipped, permanently
  • Questions the agent should refuse or escalate
  • The same question from two profiles where correct answers differ
  • Anything near column-secured fields

Correction to the ALM post. I said keep a manual list of twenty questions and run it after each deploy. Right instinct, wrong implementation — put them in a test set, grade them, run them from the pipeline.

Related: Copilot in Dynamics 365 — what it actually does


Thanks for reading this article. Hope this Article will help you. Cheers!!!

#Dynamics365 #PowerPlatform #MicrosoftCopilot #DynamicsCRM #CopilotStudio #Dataverse #AIDeveloper #DynamicsAIEngineer #HireAIEngineer #AIEngineer #D365Consultant #CRMDeveloper #AIAgents #MSDyn365

Monday, 14 September 2026

Packaging and Deploying Web Resources: A Practical ALM Workflow for Dynamics 365 JavaScript

You've written the script. You've unit-tested the logic and stepped through it in dev tools. It works perfectly on your form. Now comes the part that trips up more Dynamics 365 projects than any bug ever does: getting that script safely out of your dev environment and into production without breaking something else on the way.

This post is about the plumbing — solutions, environment variables, and a repeatable release process — rather than more JavaScript syntax. If you've been following the series, think of it as the missing link between "my script works" and "my script is live, and I can update it again next sprint without a two-hour outage window."

Web resources live in a solution — always

A web resource that exists only in your dev environment isn't deployable. The first rule of ALM in Dynamics 365 is simple: every JavaScript file, every image, every HTML web resource you build has to sit inside an unmanaged solution in dev, so it can be exported and carried forward.

Create one solution per project or product area — not the Default Solution — with your own publisher and a prefix:

Publisher: Contoso CRM Team
Prefix: con_

Web resource name: con_/scripts/opportunity.form.js
Web resource name: con_/scripts/lib/discountLogic.js

Notice the folder-style naming. Dynamics doesn't enforce folders, but the "/scripts/", "/styles/", "/images/" convention inside the name keeps a solution with 40+ web resources from turning into an unreadable flat list. Pick a convention on day one — renaming a web resource later means updating every form event handler that references it by name.

Stop hardcoding URLs and GUIDs

This is the single most common thing I see go wrong in a promotion from dev to test to production: a script with a hardcoded environment URL, flow trigger endpoint, or record GUID baked directly into the code.

// don't do this
const flowUrl = "https://prod52a1.environment.api.powerplatform.com/powerautomate/...";
const escalationTeamId = "3f9a1c20-...-...-...-prodonly";

That script works in production and breaks the moment someone imports the same managed solution into UAT, because the flow URL and team GUID are different there. Use environment variables instead — a Dataverse table type built exactly for this. Define the environment variable in your solution, give it a default value for dev, and set the current value per environment after import (test, UAT, and prod each get their own).

// do this instead
async function getEnvVarValue(schemaName) {
  const result = await Xrm.WebApi.retrieveMultipleRecords(
      "environmentvariabledefinition",
          `?$filter=schemaname eq '${schemaName}'&$expand=environmentvariabledefinition_environmentvariablevalue($select=value)`
            );
              const def = result.entities[0];
                const values = def.environmentvariabledefinition_environmentvariablevalue;
                  return values && values.length ? values[0].value : def.defaultvalue;
                  }
                  
                  const flowUrl = await getEnvVarValue("con_DiscountApprovalFlowUrl");

Yes, it's a couple more lines than a hardcoded string. It's also the difference between a 30-second solution import and a support ticket at 5pm on a Friday because someone forgot to swap a URL by hand.

A release path you can actually repeat

The environments-and-arrows diagram is familiar to anyone who's worked on a real project, but the part people skip is making it boring — the same steps, every time, ideally without a human retyping anything:

Dev (unmanaged)
  → export as MANAGED solution
      → import into Test/UAT
            → validate, set environment variable values
                    → export the same version as MANAGED
                              → import into Production
                                          → set production environment variable values

A few rules that keep this from going sideways:

Only your dev environment is unmanaged. Every downstream environment gets a managed solution. Managed solutions can't be casually edited in the target environment — which is exactly what stops someone from "just quickly" fixing a bug directly in production and having it silently overwritten by the next real release.

Bump the solution version every release (1.0.0.3, 1.0.0.4...) so you can tell at a glance what's actually deployed in each environment, and so a re-import doesn't get silently skipped as "no changes detected."

Never hand-edit a web resource inside the target environment's UI. If test or production needs a change, it goes back through dev, gets re-exported, and flows down the same path. The moment someone edits a web resource directly in production "just this once," your source of truth splits in two.

Keep the actual JavaScript in source control

The solution zip is not source control. Treat the Dataverse web resource as a deployment target, not the place your code lives. The workflow that scales:

1. Write/edit the .js file in your repo (VS Code, real linting, real diffs)
2. Use the Power Platform CLI to push it straight to your dev web resource:

   pac webresource push --file ./src/opportunity.form.js `
        --solution ConCrmTeamSolution --publisher con
        
        3. Commit the .js file to git like any other source file
        4. When it's time to release, export/pack the SOLUTION (not the .js file) via:
        
           pac solution export --name ConCrmTeamSolution --managed
              pac solution unpack --zipfile ConCrmTeamSolution.zip --folder ./solution --packagetype Managed

Once the unpacked solution folder is in git alongside your web resource source, a pull request actually shows you a meaningful diff — not a base64 blob — and code review on a form script becomes possible for the first time.

Cache is the silent killer of "but it worked in dev"

You imported the solution, the web resource clearly has your new code in the customizations, and the form still runs the old logic. Nine times out of ten, that's the browser serving a cached copy of the .js file. A few things that actually fix it, in order of how often I reach for them:

Do a hard refresh (Ctrl+Shift+R / Cmd+Shift+R) on the form before assuming your deployment failed. If it's happening to your whole team repeatedly after every release, increment a version query string convention in how the web resource is referenced, or use the "Publish All Customizations" step deliberately — a solution import doesn't always auto-publish everything, and an unpublished web resource will serve the old version even though the record itself shows the new content.

A short pre-flight checklist

Before you ship a form script anywhere past dev, it's worth running down a short list rather than trusting memory:

Is the web resource actually included in the solution (added web resources don't auto-include just because they exist)? Are there any hardcoded URLs, GUIDs, or environment-specific values left in the code? Does the target environment already have its environment variable values set, or will the script silently fail on first load after import? Did you bump the solution version? And — the one everyone forgets at least once — did you publish?

Where this leaves you

None of this is exciting the way a working Copilot Studio integration is, but it's the reason that integration keeps working three months from now, across three environments, maintained by more than one person. A script that only exists correctly in your head and your dev environment isn't shipped — it's a demo.

Next in the series, we'll go back to the form itself and look at ribbon and command bar customization with JavaScript — enable rules, custom buttons, and the display rules that decide when a command even shows up.

Copilot Studio agent ALM: getting from dev to production


Copilot Studio gets you from idea to working agent very quickly. That's the whole appeal. It also means an agent can be in front of real users before anyone's thought about how to change it safely.

Here's what I'd have in place before that happens.

Work inside a solution from the start

Agents are proper solution components, so the agent, its topics, knowledge configuration, flows, environment variables and connection references all move together as one package.

  • Create a custom publisher with your own prefix before you build anything — changing it later means recreating components
  • Keep one solution unless you genuinely need to deploy parts independently
  • Build unmanaged in dev, export managed to test and production
  • Push changes one direction only: fix in dev and redeploy, never patch production

Make anything environment-specific a variable

In dev it all just works, which is exactly why nobody notices it's hardcoded.

  • SharePoint site URLs used as knowledge sources
  • External API endpoints and base URLs
  • Notification and system email addresses
  • API keys and client secrets — use the Secret type so the value lives in Azure Key Vault, not your solution

There's a trap here worth knowing. A default value you set in dev travels with the solution. If nobody sets a proper value in the target, it quietly falls back to the dev one. That's how a production agent ends up reading your dev SharePoint site while looking perfectly healthy.

Use connection references, not connections

Credentials then bind per environment and you can run as a different account in prod. Which means deciding who the agent actually runs as — a dedicated service account or application user, with roles scoped to what it genuinely needs. Not a maker's personal account. You'll find out why the week they leave.

Promote through Pipelines

The useful part is what it checks before letting you deploy:

  • Every environment variable has a value in the target
  • Every connection reference resolves

When it fails on missing dependencies, it's usually the three dots next to the agent → Advanced → Add required objects, then try again.

The bit no tooling fixes

You can diff a plugin. You can read a pull request. You can't meaningfully diff an agent, because the change lives in instructions and topic logic — so "what changed since Tuesday" has no honest answer.

The only workaround I've found:

  • Keep twenty or so questions with known good answers
  • Cover your main topics plus whatever broke before
  • Run the lot by hand after every deploy

It's crude. It's also the only regression test you've got.

None of this is exciting work. It's just the difference between shipping version 2 and being scared to touch version 1.

Related: Copilot in Dynamics 365 — what it actually does


Thanks for reading this article. Hope this Article will help you. Cheers!!!

#Dynamics365 #PowerPlatform #MicrosoftCopilot #DynamicsCRM #CopilotStudio #Dataverse #AIDeveloper #DynamicsAIEngineer #HireAIEngineer #AIEngineer #D365Consultant #CRMDeveloper #AIAgents #MSDyn365

Display Density in Model-Driven Apps: Less Whitespace, More CRM

 If you work with Dynamics 365, you've heard this one: "Why is there so much empty space on the screen?" Microsoft now has an ...