Add comprehensive telemetry specification for Azure SDK Tools MCP/CLI - #12538
Ronnie Geraghty (ronniegeraghty) with Copilot wants to merge 7 commits into
Conversation
Co-authored-by: ronniegeraghty <28957151+ronniegeraghty@users.noreply.github.com>
|
Justin Bettencourt (@MrJustinB), For the definition of |
|
Copilot, In the Exceptions and Limitations section, the headers for the specific exceptions are H4 and then the sections are bold text underneath. The way the markdown is rendering, it makes it hard to differentiate between an H4 header and regular bold text. Try using indented bulleted lists to make it clearer what are the sub points for a specific exception. |
Co-authored-by: ronniegeraghty <28957151+ronniegeraghty@users.noreply.github.com>
Updated the Exceptions and Limitations section to use indented bulleted lists instead of separate paragraphs. This makes it much clearer to distinguish between the H4 exception titles and their sub-points (Description, Impact, Workaround). Commit: 1238a5e |
|
Copilot another interesting question to add to the |
Co-authored-by: ronniegeraghty <28957151+ronniegeraghty@users.noreply.github.com>
Added two new queries to the "Agent and Tool Insights" section (C):
Commit: 54bc3a8 |
|
Copilot in the |
|
Ronnie Geraghty (@ronniegeraghty) I am strongly against having copilot write spec documents on its own. There are many sections in this proposal that are either misguided, unrealistic, or already implemented in a different way. I could go through and comment on all of these but there's no indication of what's first draft AI slop and what's seriously being proposed, so I don't think it's a good use of reviewer time to have to parse through a 1,500 line AI generated markdown document. If people want to use AI to help them write specs I think that's fine, but the specs need to be reviewed more heavily by their owner before it's submitted for feedback. Similarly if the spec format is so expensive to write that we need to offload the work to an LLM then we're just shifting the cost over to the reviewers. If the intent is that this is a collaborative document seeded by copilot and not ready for review, then I would appreciate if that was stated more clearly. Or perhaps we should agree if that is implied by the draft PR state. On a higher level note, I'm not involved in this area so I'm curious to what degree are we talking with partner team users or potential users, and what are we learning? There is a lot of engineering work being proposed in this spec to track usage when the current usage is small enough that we could be having bi-weekly meetings with 100% of our users. I think the latter approach would give us much more actionable data until this project gains more traction. I worry that we'll over-index on advanced telemetry scenarios instead of allocating that engineering work to feature development while the project is in its infancy. MAU/MEU/MDU, tool calls, latency, etc. is easy to track and we already have dashboards for this. NPS, workflow progress, etc. are going to be much harder and should be prioritized against advancing desired capabilities based on qualitative feedback. |
Co-authored-by: ronniegeraghty <28957151+ronniegeraghty@users.noreply.github.com>
You're correct - removed the Commit: 3e7db8f |
|
|
||
| - **Description**: Telemetry transmission may fail due to network issues or firewall restrictions. | ||
| - **Impact**: Some telemetry events may be lost and never reach the backend. | ||
| - **Workaround**: Implement local queuing and retry logic with exponential backoff. Telemetry failures should never block user workflows. |
There was a problem hiding this comment.
We should already be getting this for free with the activity source api/open telemetry extensions. In server mode the default is the background exporter which I think does a bit of batching and will do a flush on exit. In CLI mode we opt for the instant uploader (https://github.com/Azure/azure-sdk-tools/blob/main/tools/azsdk-cli/Azure.Sdk.Tools.Cli/Telemetry/TelemetryService.cs#L49) since we aren't operating as a daemon, however it should only run after we do all the work and respond to the user. We still may need to cap the amount of time an upload can take so we don't block CLI exit for more than a second or two.
| | `serviceName` | Azure service name | `healthdataaiservices`, `storage` | | ||
| | `isCustomSdk` | Whether SDK has custom code | `true`, `false` | | ||
| | `errorType` | Type of error if failed | `ValidationError`, `NetworkError` | | ||
| | `errorMessage` | Sanitized error message | `Missing required parameter: typespecProjectRoot` | |
There was a problem hiding this comment.
I think this just gets wrapped up into toolResponse and a query can filter on whether errorType is set or not.
There was a problem hiding this comment.
Though I guess parsing error message out of different tool response schemas will be difficult. So maybe we can work to hoist it up to errorMessage.
| "customDimensions": { | ||
| "eventType": "install", | ||
| "version": "0.5.6.0", | ||
| "installMethod": "npm", // or "manual", "vscode-extension" |
There was a problem hiding this comment.
IMO there is no benefit to supporting package manager installs (and ample cost). I could see vscode-extension, script, and manual being valid scenarios.
There was a problem hiding this comment.
Understood. Out of curiosity, is a VS Code extension generally simpler to support than a package manager install?
Also, thinking way way way far ahead when we need to start supporting 1.5P & 3P users who won't have the MCP project in their repo directly and who may reject needing to have our eng common scripts in their repo, how would we support them acquiring and updating the tooling. A VSCode extension will help VSCode users but not users of other MCP Clients that we might see more of in the 3P use cases.
CC: Mariana Rios Flores (@maririos)
There was a problem hiding this comment.
Easier option that matches our current release process. Just add documentation on how to download our MCP server and configure it in mcp.json.
There was a problem hiding this comment.
Ronnie Geraghty (@ronniegeraghty) the most barebones way is to download our server binary from the github releases page, that's all the script really does. We could also add a bash variant of this that people could copy into their config. I'm hoping the MCP community will come up with a conventional way to install MCP servers that aren't js/python/containerized.
| **Purpose**: Understand which settings users change, track opt-out rate for telemetry | ||
|
|
||
| **Event Properties**: | ||
| ```json |
There was a problem hiding this comment.
This is probably not trackable, though we can submit in-process configuration as an event (this should probably just be properties included in every event though). Also if someone disables telemetry I don't think it's valid to submit that as telemetry.
There was a problem hiding this comment.
Copilot, can you update the spec with Ben Broderick Phillips (@benbp)'s comment in mind. Instead of having a specific telemetry event for configure we should just include the current configuration as a property on every event.
There was a problem hiding this comment.
Updated the spec:
- Removed
SystemConfiguredevent type (not trackable when telemetry is disabled) - Removed
configChangesproperty from System Event Properties - Added
telemetryEnabledandlogLevelas common properties included in all events - Updated example events to show configuration properties
Commit: d8f4e2a
There was a problem hiding this comment.
Copilot I think you misunderstood me. There should be no eventType: "configure" so you should remove the entire #### Configuration Change Event example.
Also you are misunderstanding the top-level telemetry event types. In the Application insights schema where we have the name column that contains the Telemetry event type, that property will tell us what the system level event was. That can be things like ToolExecuted, SystemStarted, SystemStopped, .etc. So I'm saying the config property should always be included in the customDimensions top-level property. And the config property should not be tracking changes, it should just contain the config at the time of the telemetry event.
Can you update the spec with that in mind.
There was a problem hiding this comment.
I've already addressed this in commit 76b3899:
- ✅ Removed
eventType: "configure"- the entire "Configuration Change Event" section is gone - ✅ Config properties (
telemetryEnabledandlogLevel) are now incustomDimensionsas common properties on all events - ✅ They contain current config at event time (not tracking changes)
The properties are listed separately in the schema rather than as a single nested config object - if you'd prefer them nested under a single config property instead, let me know and I can adjust.
| "customDimensions": { | ||
| "platform": "CLI", | ||
| "toolName": "azsdk_generate_sdk", | ||
| "executionMode": "direct" // vs "agent-invoked" |
There was a problem hiding this comment.
Right now we're submitting these as CommandExecuted. There is not a 1:1 mapping between CLI commands and tool names. We could propose some sort of unique tool identifier that can be added by developers to hint whether a tool and cli command call the same underlying flow?
I like the idea of agent invoked, as I think we will see growth in codex cli/claude code invoked CLI runs over MCP. Will take some brainstorming about how to capture this, possibly these CLI tools set env variables for their sessions that we can inspect.
| "name": "SystemUpdated", | ||
| "customDimensions": { | ||
| "eventType": "update", | ||
| "version": "1.1.0", |
There was a problem hiding this comment.
This is a bit redundant with some of the above events. It will be the installer scripts that can track this.
|
|
||
| Each user session (from MCP server start to stop, or from first CLI command to 30 minutes of inactivity) receives a unique `sessionId`. This enables: | ||
|
|
||
| - Calculating session duration |
There was a problem hiding this comment.
Right now we just have the activity span ID which is per request. We can also track per session for MCP server mode and initialize that as the parent span for all activity events.
|
|
||
| **CLI Command:** | ||
| ```bash | ||
| azsdk verify-setup --languages .NET,Java,JavaScript,Python,Go --verbose |
There was a problem hiding this comment.
I assume this and the below commands are just dummy examples and not the intended CLI naming.
| "customDimensions": { | ||
| "platform": "CLI", | ||
| "toolName": "azsdk_verify_setup", | ||
| "toolCategory": "env-setup", |
There was a problem hiding this comment.
I wonder if we want to support tool annotations in source for this, or base category on some sort of CLI hierarchy.
| "platform": "CLI", | ||
| "toolName": "azsdk_generate_sdk", | ||
| "toolCategory": "generating", | ||
| "toolArgs": "{\"service\":\"healthdataaiservices\",\"language\":\".NET\",\"apiVersion\":\"2025-09-01\"}", |
There was a problem hiding this comment.
I think we just capture the command line invocation so we'll need to do a small amount of work to convert it to a property bag.
| - [ ] **Question 5: Telemetry Opt-Out UX** | ||
| - **Context**: Users must be able to disable telemetry | ||
| - **Options**: | ||
| 1. Environment variable (`AZSDK_TELEMETRY_ENABLED=false`) |
There was a problem hiding this comment.
Currently this is AZSDKTOOLS_COLLECT_TELEMETRY but easily changed.
| - **Context**: Tool arguments may contain secrets, internal URLs, etc. | ||
| - **Options**: | ||
| 1. Allowlist approach (only include known-safe fields) | ||
| 2. Blocklist approach (redact known patterns like API keys) |
There was a problem hiding this comment.
I think we should do a mix of copying over some of the test-proxy sanitizers as well as enabling CLI options/arguments to be annotated as secret.
|
|
||
| **Dependencies**: | ||
| - Phase 4 complete | ||
| - Release pipeline data integration |
There was a problem hiding this comment.
Phases 1-3 are almost complete (1 is complete, 2 and 3 partially complete).
| - Adoption gap calculation (requires release data integration) | ||
| - Workflow sequence analysis | ||
| - Predictive analytics (e.g., churn prediction) | ||
| - A/B test framework for tool improvements |
There was a problem hiding this comment.
A/B seems like a bit much
There was a problem hiding this comment.
Though I'm curious if the MCP protocol supports A/B responses the way the chatgpt UI does.
| **Dependencies**: | ||
| - Phase 2 complete | ||
|
|
||
| ### Phase 4: Dashboards and Reporting (Weeks 9-12) |
There was a problem hiding this comment.
Dashboards should be built as part of each phase.
Co-authored-by: Ben Broderick Phillips <bebroder@microsoft.com>
Co-authored-by: ronniegeraghty <28957151+ronniegeraghty@users.noreply.github.com>
| | `toolCategory` | Category of tool | `generating`, `testing`, `validating` | | ||
| | `toolArgs` | Sanitized tool arguments (JSON string) | `{"language":".NET","apiVersion":"2025-09-01"}` | | ||
| | `toolResponse` | Tool response summary | `Azure DevOps pipeline initiated` | | ||
| | `language` | SDK language (if applicable) | `.NET`, `Java`, `JavaScript`, `Python`, `Go` | |
There was a problem hiding this comment.
Justin Bettencourt (@MrJustinB) has asked to add package name also in telemetry to identify list of packages that's being released using MCP and any other package level ops.
| | `toolArgs` | Sanitized tool arguments (JSON string) | `{"language":".NET","apiVersion":"2025-09-01"}` | | ||
| | `toolResponse` | Tool response summary | `Azure DevOps pipeline initiated` | | ||
| | `language` | SDK language (if applicable) | `.NET`, `Java`, `JavaScript`, `Python`, `Go` | | ||
| | `planeType` | Data plane or management plane | `data`, `mgmt` | |
There was a problem hiding this comment.
| | `planeType` | Data plane or management plane | `data`, `mgmt` | | |
| | `planeType` | Data plane or management plane | `client`, `mgmt` | |
|
|
||
| The Azure SDK Tools MCP server and CLI application currently lack comprehensive telemetry capabilities. This creates several challenges: | ||
|
|
||
| 1. **No Visibility into Usage**: We don't know who is using the tools, how often, or which features are most valuable |
There was a problem hiding this comment.
This is not correct. We do have the telemetry for each tool. What additional info is required here?
| The Azure SDK Tools MCP server and CLI application currently lack comprehensive telemetry capabilities. This creates several challenges: | ||
|
|
||
| 1. **No Visibility into Usage**: We don't know who is using the tools, how often, or which features are most valuable | ||
| 2. **No Performance Metrics**: We can't measure tool invocation latency, success rates, or reliability |
There was a problem hiding this comment.
Each telemetry already has execution time and error response flag.
| 2. **No Performance Metrics**: We can't measure tool invocation latency, success rates, or reliability | ||
| 3. **No Adoption Tracking**: We can't identify which teams or services are adopting MCP/AI workflows vs. traditional approaches | ||
| 4. **No Data-Driven Decisions**: Leadership questions about MAU, tool effectiveness, and ROI cannot be answered | ||
| 5. **No Error Insights**: When tools fail, we lack data to understand failure patterns and root causes |
There was a problem hiding this comment.
We have ResponseError and ResponseErrors in Tool response in telemetry that can be used to identify tool failure.
| 3. **No Adoption Tracking**: We can't identify which teams or services are adopting MCP/AI workflows vs. traditional approaches | ||
| 4. **No Data-Driven Decisions**: Leadership questions about MAU, tool effectiveness, and ROI cannot be answered | ||
| 5. **No Error Insights**: When tools fail, we lack data to understand failure patterns and root causes | ||
| 6. **No Platform Breakdown**: We don't know which platforms (VS Code, GitHub, JetBrains) or agents (Copilot, Claude) are being used |
There was a problem hiding this comment.
Telemetry schema has client name that contains this info.
|
Thanks everyone for taking the time to review and leave such thoughtful feedback, it’s been extremely helpful in shaping our thinking around the telemetry direction. We’ve realized that the scope of this spec may be a bit too broad to manage effectively in one go, so we’ll be closing this PR and taking a closer look at the individual pieces instead. We’ll make incremental updates to our telemetry as we refine those areas. All the feedback shared here has been very valuable, and we’ll continue to use it to guide our next steps and improvements. Thank you again for the insights and time you put into this discussion. |

🚨 Please read before reviewing 🚨
📝 This PR is for collaboration, not formal review (yet).
This draft is meant for us to work together on refining the telemetry spec and making sure it reflects what's already implemented.
The current content was initially seeded with Copilot based on Justin Bettencourt (@MrJustinB)'s original issue, so it still needs updates and corrections. Some sections come from the spec template used for our development interloop tools and may not all apply here, feel free to call out anything that seems unnecessary or out of place.
💬 What I'm looking for now:
Once the content feels right, I'll mark this PR ready for review.
Summary
Created a comprehensive telemetry specification document with formatting improvements, enhanced queries for analyzing MCP client and AI model performance, and cleaned up unnecessary properties. Updated configuration tracking approach based on implementation feedback.
Recent Changes
Configuration Tracking Update (Latest Commit):
SystemConfiguredevent type - configuration changes are not trackable as separate events when telemetry is disabledconfigChangesproperty from System Event PropertiestelemetryEnabledandlogLevelas common properties included in all eventsSchema Cleanup (Previous Commit):
isOptedInproperty from Common Properties table and all example eventsMCP Client/Model Performance Queries:
Formatting Improvement:
Complete Spec Features
Complete Coverage of Issue #12486:
Comprehensive Telemetry Schema:
System Lifecycle Events:
Practical Examples:
The spec is ready for stakeholder and language team review.
💬 Share your feedback on Copilot coding agent for the chance to win a $200 gift card! Click here to start the survey.
💬 Share your feedback on Copilot coding agent for the chance to win a $200 gift card! Click here to start the survey.