# (RFC)Architecture Design: Scalable OpenTelemetry for Jenkins

**URL:** <https://community.jenkins.io/t/rfc-architecture-design-scalable-opentelemetry-for-jenkins/36060>\
**Category:** GSoC\
**Created:** [January 16, 2026, 8:16pm UTC](https://community.jenkins.io/t/rfc-architecture-design-scalable-opentelemetry-for-jenkins/36060 "2026-01-16T20:16:08Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Zenith1415](https://dub1.discourse-cdn.com/flex013/user_avatar/community.jenkins.io/zenith1415/32/19194_2.png) [@Zenith1415](https://community.jenkins.io/u/Zenith1415)\
**Post date:** [January 16, 2026, 8:16pm UTC](https://community.jenkins.io/t/rfc-architecture-design-scalable-opentelemetry-for-jenkins/36060/1 "2026-01-16T20:16:08Z")

</div>

Hi everyone,

I am exploring the ‘Use OpenTelemetry for Jenkins Jobs’ project for GSoC 2026. Following a discussion with maintainers on Gitter, I am opening this thread to document the architectural design and trade-offs for ci.jenkins.io.The Challenge Enabling the OTel plugin on the Jenkins infrastructure scale is not just about installation, it requires an engineering strategy to handle High Cardinality and Storage Costs.

Proposed Architecture I have been researching a Backend-Agnostic OTLP Pipeline where the Collector acts as a gateway. The key feature is Tail-based Sampling, which allows us to:

- Retain 100% of failed build traces (for debugging).

- Aggressively sample successful builds (to save storage).

- Switch backends (Jaeger, Tempo, Prometheus) without changing Jenkins configuration.

Proof of Concept & Trade-off Analysis I have set up a local lab (Jenkins → Collector → Jaeger) to validate the probabilistic\_sampler and load generation. I have compiled my findings, including a detailed analysis of Cardinality vs. Query Speed, in the draft [(RFC)Architecture Design: Scalable OpenTelemetry for Jenkins - Google Docs](https://docs.google.com/document/d/13uJDFPIZjFVSX1L95qIznFgvOofEcXyCNC7Exzjlcr0/edit?usp=sharing)

I would appreciate any feedback from the Infra team, specifically regarding the retention policies for successful builds.

Amanraz Thakur

---

<div class="post-metadata">

**Author:** ![Zenith1415](https://dub1.discourse-cdn.com/flex013/user_avatar/community.jenkins.io/zenith1415/32/19194_2.png) [@Zenith1415](https://community.jenkins.io/u/Zenith1415)\
**Post date:** [February 7, 2026, 11:46am UTC](https://community.jenkins.io/t/rfc-architecture-design-scalable-opentelemetry-for-jenkins/36060/2 "2026-02-07T11:46:13Z")

</div>

**Update on the Context Propagation Fix**

**Update:** I have successfully implemented and validated the fix for the `TRACEPARENT` context issue.

**The Issue:** Previously, the `TRACEPARENT` environment variable was not correctly updating for steps inside a stage. This caused broken trace hierarchies where steps like `sh` or `bat` were not properly nested under their parent stage.

**The Fix:** I modified `OtelEnvironmentContributor.java` to explicitly respect the current active span. You can view the implementation here: **[[Link to The PR]](https://github.com/jenkinsci/opentelemetry-plugin/pull/1219)**

**Visual Verification:** I spun up a local Jenkins instance connected to a Jaeger backend to verify the fix. As seen in the screenshot below, the `sh` step is now correctly indented as a child of the `Verification` stage, confirming that the trace context is being propagated correctly down the pipeline.

 ![trace_fix_proof](https://europe1.discourse-cdn.com/flex013/uploads/jenkins/original/3X/3/b/3b46de788fc5785e537e4d4061e37e0eb02a4b57.png)
