Skip to content

Comments

.NET: Fixing issue where OpenTelemetry span is never exported in .NET in-process workflow execution#4196

Open
alliscode wants to merge 6 commits intomicrosoft:mainfrom
alliscode:investigation/issue-4155-1771877003
Open

.NET: Fixing issue where OpenTelemetry span is never exported in .NET in-process workflow execution#4196
alliscode wants to merge 6 commits intomicrosoft:mainfrom
alliscode:investigation/issue-4155-1771877003

Conversation

@alliscode
Copy link
Member

@alliscode alliscode commented Feb 23, 2026

This pull request addresses the issue where workflow run telemetry spans (Activity objects) were not always properly stopped and exported, particularly in streaming and lockstep execution environments. The changes ensure that workflow run activities are disposed as soon as the workflow reaches the idle state or when the run loop exits, preventing telemetry data from being lost. Additionally, comprehensive regression tests are added to verify correct activity lifecycle management.

Improvements to Activity Lifecycle Management:

  • Ensured that the workflow.run Activity is disposed immediately when the workflow reaches the Idle state, so telemetry spans are promptly exported rather than waiting for cancellation or disposal.
  • Added a safety net to dispose of the workflow.run Activity if it was not already stopped when the run loop exits, covering cancellation and error scenarios.
  • Removed the using statement from the activity initialization to allow manual control over the activity's disposal timing.

Testing and Regression Coverage:

  • Added a new test file WorkflowRunActivityStopTests.cs to verify that workflow run activities are always properly stopped and exported to telemetry backends, covering lockstep, off-thread, and streaming execution environments, as well as ensuring that all started activities are stopped.

Closes #4155

Contribution Checklist

  • The code builds clean without any errors or warnings
  • The PR follows the Contribution Guidelines
  • All unit tests pass, and I have added new tests where possible
  • Is this a breaking change? If yes, add "[BREAKING]" prefix to the title of the PR.

…ity never stopped in streaming OffThread path

The WorkflowRunActivity_IsStopped_Streaming_OffThread test demonstrates that
the workflow.run OpenTelemetry Activity created in StreamingRunEventStream.RunLoopAsync
is started but never stopped when using the OffThread/Default streaming execution.
The background run loop keeps running after event consumption completes, so the
using Activity? declaration never disposes until explicit StopAsync() is called.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

2. Fix workflow.run Activity never stopped in streaming OffThread execution (microsoft#4155)

The workflow.run OpenTelemetry Activity in StreamingRunEventStream.RunLoopAsync
was scoped to the method lifetime via 'using'. Since the run loop only exits on
cancellation, the Activity was never stopped/exported until explicit disposal.

Fix: Remove 'using' and explicitly dispose the Activity when the workflow reaches
Idle status (all supersteps complete). A safety-net disposal in the finally block
handles cancellation and error paths.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot AI review requested due to automatic review settings February 23, 2026 22:10
@markwallace-microsoft markwallace-microsoft added .NET workflows Related to Workflows in agent-framework labels Feb 23, 2026
@github-actions github-actions bot changed the title Fixing issue where OpenTelemetry span is never exported in .NET in-process workflow execution .NET: Fixing issue where OpenTelemetry span is never exported in .NET in-process workflow execution Feb 23, 2026
Copy link
Contributor

Copilot AI left a comment

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR aims to ensure OpenTelemetry workflow-run spans (Activity) are reliably stopped/disposed (and therefore exported) during .NET in-process workflow execution, including streaming scenarios, and adds regression tests around activity lifecycle behavior.

Changes:

  • Updated StreamingRunEventStream.RunLoopAsync to manually manage the workflow-run Activity lifecycle (stop on Idle and ensure disposal on loop exit).
  • Added WorkflowRunActivityStopTests to assert workflow-run activities are started and stopped across multiple execution modes.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 3 comments.

File Description
dotnet/src/Microsoft.Agents.AI.Workflows/Execution/StreamingRunEventStream.cs Changes workflow-run Activity disposal timing to stop/export spans earlier and adds a safety-net disposal on exit.
dotnet/tests/Microsoft.Agents.AI.Workflows.UnitTests/WorkflowRunActivityStopTests.cs Adds regression coverage validating workflow-run activities are stopped/disposed in lockstep, off-thread, and streaming usage.

…\nImplements two-level telemetry hierarchy per PR feedback from lokitoth:\n- workflow.session: spans the entire run loop / stream lifetime\n- workflow_invoke: per input-to-halt cycle, nested within the session\n\nThis ensures the session activity stays open across multiple turns,\nwhile individual run activities are created and disposed per cycle.\n\nAlso fixes linkedSource CancellationTokenSource disposal leak in\nStreamingRunEventStream (added using declaration)."
Copy link
Contributor

Copilot AI left a comment

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 7 out of 7 changed files in this pull request and generated 7 comments.

@alliscode
Copy link
Member Author

@copilot open a new pull request to apply changes based on the comments in this thread

@alliscode
Copy link
Member Author

@copilot open a new pull request to apply changes based on the comments in this thread

…dd error tag\n\n1. LockstepRunEventStream: Remove 'using' from Activity in async iterator\n and manually dispose in finally block (fixes microsoft#4155 pattern). Also dispose\n   linkedSource CTS in finally to prevent leak.\n2. Tags.cs: Add ErrorMessage (\"error.message\") tag for runtime errors,\n   distinct from BuildErrorMessage (\"build.error.message\").\n3. ActivityNames: Rename WorkflowRun from \"workflow_invoke\" to \"workflow.run\"\n   for cross-language consistency.\n4. WorkflowTelemetryContext: Fix XML doc to say \"outer/parent span\" instead\n   of \"root-level span\".\n5. ObservabilityTests: Assert WorkflowSession absence when DisableWorkflowRun\n   is true.\n6. WorkflowRunActivityStopTests: Fix streaming test race by disposing\n   StreamingRun before asserting activities are stopped.\n7. StreamingRunEventStream/LockstepRunEventStream: Use Tags.ErrorMessage\n   instead of Tags.BuildErrorMessage for runtime error events."
…urce, move SessionStarted earlier\n\n- Revert ActivityNames.WorkflowRun back to \"workflow_invoke\" (OTEL semantic convention contract)\n- Use 'using' declaration for linkedSource CTS in LockstepRunEventStream (no timing sensitivity)\n- Move SessionStarted event before WaitForInputAsync in StreamingRunEventStream to match Lockstep behavior"
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

.NET workflows Related to Workflows in agent-framework

Projects

None yet

Development

Successfully merging this pull request may close these issues.

.NET: [Bug]: workflow.run OpenTelemetry span is never exported in .NET in-process workflow execution

4 participants