OpenTelemetry Is Stuck: A Spreadsheet Shows Why

OTel isn't going well (and I made a spreadsheet about it)

OpenTelemetry Is Stuck: A Spreadsheet Shows Why

OpenTelemetry, the vendor-neutral observability standard, is struggling. A data-driven analysis of CNCF projects reveals that most language SDKs rely on just one or two maintainers, while the semantic-conventions repo drags on with PRs taking months. The root cause: a binary stability gate that makes every feature a permanent commitment, combined with a massive scope and too few paid maintainers. The result is a project frozen by its own caution, where progress is slow and burnout is real.

The issue is more a classic case of "someone has to pay the maintainers".
  1. osener

    I like the end result of OpenTelemetry tracing when using Axiom and the like, but the SDKs have been a nightmare. Too much emphasis on automatic instrumentation, Java-isms, everything is stateful and abstracted away.

    It can do distributed tracing of otherwise traditional long running microservices, but breaks down when your functions are distributed like in durable execution engines, Cloudflare Workflows, “functions” that span hours/days/weeks and steps that retry many times.

    I had to reverse engineer how SDKs work and how tracing UIs display data so I could make simpler functions that fit wider variety of runtimes and more freely parent spans, start spans and end them from different function instances.

    I think most of the API and terminology complexity is self inflicted. Would love to see a rebooted developer experience that is less Kubernates-brained.

  2. EdSchouten

    What always puzzles me about OpenTelemetry is that tracing, metrics and logs are all designed independently. I wish there was a way I could just annotate my code base once, and let the ultimate decision to expose something as a metric/log/trace be dynamic at runtime.

    For example, if I look at a graph in monitoring dashboard and see something suspicious, I’d like to say: “The next time something like this occurs again, please save me a trace.” I should be able to just do that with a single mouse click.

    I remember them releasing the tracing spec/SDKs and saying “now let’s move on to metrics/logs.” That never sat right with me.

  3. dijit

    I know sadly very little about otel, it feels “heavy” in a way I am not used to, I am used to simple systems - configured and composed in a way that makes a larger system.

    20 years ago, we were doing (what I think) OTel is doing: with “hit IDs” (half way between a session and a request) that were consistently applied when logging the cause a request being fired; along centralised logging and really good timekeeping. Essentially a unique identifier as a tag that followed the request as it passed through the system.

    This was enough to debug basically any problem.

    We could even measure the distance between requests of the same “hit” and the total wall-time before it managed to return through the load balancer, so we could track our p99 easily.

    Though truthfully we didn't make pretty graphs.

    I sometimes wonder what OTel gives me more than this, but I work in games now and lots of these things that work well in webdev do not apply at all to our problems.

  4. czhu12

    It really never grokked with me why there isn't just "open source Datadog" that can be installed and used. End to end, stateful, that we can just self host.

    Our team tried to set up open telemetry to replace Datadog and got totally crushed in complexity. The model of having Open Telemetry just be for standardizing & exporting to other backends, needing glue for each part of the setup was nuts.

  5. Havoc

    I find the entire observability space to quite a poor experience, at least in the self-hosted space. Tried both grafana route and signoz and neither seems particularly pleasant

  6. brikym

    I've never found instrumentation to be a huge issue. Sure it takes more effort but you get a lot more value once you understand _business_ events.

  7. tete

    OpenTelemtry is the perfect example of an overengineered mess.

    While I usually think that at least having some standard that people agree on I think OpenTelemtry should be dropped.

    A lot of the less popular alternatives (just going with Prometheus, Victoriametrics, etc) are de-facto competing smaller standards and a lot better both in terms of less added complexity and the results you get.

    I think OpenTelemetry turned metrics into a farce. In many situations even self-rolled telemetry works better even with the added stuff. The annoying thing is that OpenTelemtry is that big standard now one kind of has to to add compatibility. So please, if you write software, make sure you don't lock yourself into OTel.

  8. rcleveng

    Sounds a lot like K8s. It's not a framework you use, it's a framework to build a framework on top of.

    I wish the observability vendors would move to using it under the covers so it's easier to mix and match.

    I wish the otel support wasn't super buggy in most of the frameworks and backends.

More from this day

2026-08-22