Open Source Archives | Dynatrace news https://www.dynatrace.com/news/category/open-source/ The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Fri, 12 Jun 2026 09:12:46 +0000 en hourly 1 Adriana Villela of Dynatrace takes on OpenTelemetry community manager role https://www.dynatrace.com/news/blog/adriana-villela-of-dynatrace-takes-on-opentelemetry-community-manager-role/ https://www.dynatrace.com/news/blog/adriana-villela-of-dynatrace-takes-on-opentelemetry-community-manager-role/#respond Mon, 02 Feb 2026 20:53:14 +0000 https://www.dynatrace.com/news/?p=72965 Dynatrace and OpenTelemetry

Dynatrace principal developer advocate Adriana Villela is stepping up her involvement with the OpenTelemetry® community. As of today, she will be taking on the role of OpenTelemetry community manager, alongside Reese Lee of New Relic. They are joined by Julia Furst Morgado, of Dash0, who will serve as associate community manager, and will be taking […]

The post Adriana Villela of Dynatrace takes on OpenTelemetry community manager role appeared first on Dynatrace news.

]]>
Dynatrace and OpenTelemetry

Dynatrace principal developer advocate Adriana Villela is stepping up her involvement with the OpenTelemetry® community. As of today, she will be taking on the role of OpenTelemetry community manager, alongside Reese Lee of New Relic. They are joined by Julia Furst Morgado, of Dash0, who will serve as associate community manager, and will be taking over the role from Austin Parker of Honeycomb.

Community managers are appointed by the OpenTelemetry Governance Committee and act as public stewards of the contributor and end-user community. They’re responsible for organizing events, coordinating efforts to improve contributor experience, managing the OpenTelemetry presence on social media, and working to nurture and grow the OpenTelemetry community overall. Effectively, this is a project-maintainer role whose project is the OpenTelemetry community itself.

“I’m thrilled, then, that Adriana and Reese have stepped up to not only continue growing our community management role, but invest time in helping scope and shape this function for the project as it enters into its next era of growth,” said Austin in the official announcement, which was also made at OTel Unplugged in Brussels, Belgium, on February 2nd, 2026.

“As a community manager, I want to continue to raise the OpenTelemetry community profile within the Cloud Native Computing Foundation and the wider open-source world,” Adriana says.

The path to OpenTelemetry community leadership

Adriana’s journey with the OpenTelemetry community began in 2021 when she worked for Tucows, where she managed the platform engineering and observability teams. “I had dabbled a bit in observability when I was at my previous job as a release engineer at Ceridian, but when I stepped in as manager of the observability team at Tucows, to do right by my team and the organization, so I could lead them in the right direction,” she says. “I learned as much as I could about observability and OpenTelemetry projects, and I did it in public.” She documented her learnings in her Unpacking Observability series on Medium.

While at Tucows, Adriana made it her mission to move away from vendor lock-in and towards OpenTelemetry tools. But 2021 was still early days for the project. For example, traces weren’t yet generally available. To learn more about how to expand OpenTelemetry project use in a large enterprise, she connected with OpenTelemetry project co-founder Ted Young, then director of open source development at Lightstep, and Honeycomb field CTO Liz Fong-Jones, then developer advocate at Honeycomb. “Both Liz and Ted put aside their competitor differences to jump on a Q&A call with Tucows developers to answer their questions and concerns about OTel,” Adriana recalls.

Adriana’s writings caught Austin’s eye in 2022, and they hired her for her first developer relations role. As part of the role, she was encouraged to contribute to the OpenTelemetry community. When one of the original founders of the OpenTelemetry End User Working Group (later converted to a special interest group or SIG) left the project in 2023, Ted and Austin asked Adriana to help lead the group alongside Reese.

Together, Adriana and Reese raised the End User SIG’s profile by running monthly Open Telemetry livestreams, launched the Humans of Open Telemetry series, and collaborated with various other SIGs to gather end-user feedback through surveys. She also contributed to the CNCF OpenTelemetry Certified Associate (OTCA) certification.

The OpenTelemetry community has advanced significantly since Adriana began her exploration of the project, and it only continues to evolve. She points to the work around the Open Agent Management Protocol (OpAMP) for managing data collection agents such as the OpenTelemetry Collector, and the increased emphasis on the quality of telemetry data. Over the years, most of the major observability vendors have embraced OpenTelemetry tooling, both by supporting its data ingest format, OTLP, natively, and as by contributing to OpenTelemetry projects, making it the standard for instrumentation. “That tells me that it’s here to stay,” she says. “I can’t wait to see where the community takes it next.”

Reflections from the OpenTelemetry community on Villela’s impact

Comments from Adriana Villela’s fellow contributors on her engagement with the OTel community:

“Adriana is a tireless and omnipresent anchor of the OpenTelemetry community, whether it’s educating, building, or the organizing the end user SIG. My entry into the community was easy thanks to her efforts, and I’m sure countless others could say the same.”

— Josh Lee, Open Source Evangelist at Altinity

“I met Adriana at the very first OTel Unplugged in 2022, and together we developed the End User Working Group from a 3-person task force to the bigger and better End User SIG it is today (we’ve since added 2 more Maintainers!). I’m so stoked to be moving to the next stage of OTel community work alongside Adriana.”

— Reese Lee, Developer Relations Engineer and incoming OTel Community Manager, New Relic

“Adriana has a unique voice in the OpenTelemetry community, addressing current both end user and community challenges. Through her technical work and her public contributions in the open source observability ecosystem, Adriana connects with a wider audience: from developers to platform engineers or first time contributors. Her blog posts, conference talks and podcasts, help practitioners understand not just how OpenTelemetry works, but why it works the way it does, creating a trail of useful resources that anybody can use for onboarding or deepening specific topics.”

— Diana Todea, Developer Experience Engineer and OpenTelemetry contributor, Victoria Metrics

Dynatrace teams support over 30 open source projects and have made 60k+ commits to OpenTelemetry projects. Learn how you can get the most out of your telemetry data with Dynatrace.

The post Adriana Villela of Dynatrace takes on OpenTelemetry community manager role appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/adriana-villela-of-dynatrace-takes-on-opentelemetry-community-manager-role/feed/ 0
What is OpenTelemetry?  An open-source standard for logs, metrics, and traces https://www.dynatrace.com/news/blog/what-is-opentelemetry/ https://www.dynatrace.com/news/blog/what-is-opentelemetry/#respond Tue, 15 Jul 2025 14:43:50 +0000 https://www.dynatrace.com/news/?p=69968 OpenTelemetry and Dynatrace make a winning combination

OpenTelemetry is an open-source framework of tools, APIs, and SDKs that help analysts understand software performance and behavior. Also referred to as OTel, OpenTelemetry is rapidly solidifying its position as a fundamental tool in the world of observability. Born as an open-source project under the Cloud Native Computing Foundation (CNCF), OpenTelemetry provides a unified framework […]

The post What is OpenTelemetry?  An open-source standard for logs, metrics, and traces appeared first on Dynatrace news.

]]>
OpenTelemetry and Dynatrace make a winning combination


OpenTelemetry is an open-source framework of tools, APIs, and SDKs that help analysts understand software performance and behavior. Also referred to as OTel, OpenTelemetry is rapidly solidifying its position as a fundamental tool in the world of observability.

Born as an open-source project under the Cloud Native Computing Foundation (CNCF), OpenTelemetry provides a unified framework for generating, collecting, processing, and exporting telemetry data—including logs, metrics, and traces.

Using OpenTelemetry,  IT teams can instrument, generate, collect, and export telemetry data for analysis to better understand software performance and behavior. When OpenTelemetry debuted in beta in 2020, it replaced its predecessors, OpenTracing and OpenCensus.

OpenTelemetry enables observability

To appreciate what OTel does, it helps to understand observability. Traditionally speaking, observability is the ability to understand what’s happening inside a system from the knowledge of the external data it produces; usually logs, metrics, and traces.

But the data itself is only as good as what you can learn from and do with it. This definition from Hazel Weakly sums it up nicely:

“Observability is the ability to ask meaningful questions, get useful answers, and act effectively on what you’ve learned.”

Observability is important because the systems of today are exponentially more complex than the systems of ten, or even five years ago. The shift from monolithic to distributed IT architectures introduces many more moving parts and interactions to keep track of, sometimes leading to systems behaving in unpredictable ways. Observability helps you make sense of what’s happening so you can act on this information, and OpenTelemetry helps to enable observability.

By promoting consistency and interoperability, OpenTelemetry enhances observability practices and benefits the entire industry by streamlining and standardizing how everyone can collect and use data.

Since the project’s start, many vendors, including Dynatrace, have contributed to the project to make rich data collection easier and more consumable. In fact, Dynatrace is one of the top contributing organizations to OpenTelemetry.

Benefits of OpenTelemetry

Collecting application data is nothing new. However, the collection mechanism and format are rarely consistent from one application to another. This inconsistency can be a nightmare for developers and Site Reliability Engineers (SREs) who are just trying to understand the health of an application.

Most of the major observability vendors, including Dynatrace, support OTel. As a result, it has become the de facto standard for instrumenting cloud-native applications. What differentiates observability solutions from one another is what they do with your data to help you ask the right questions. Asking the right questions unlocks an elevated level of understanding, giving businesses the ability to accelerate growth, drive innovation, and deliver experiences customers love.

It’s akin to how Kubernetes became the standard for container orchestration. This broad adoption has made it easier for organizations to implement container deployments since they don’t need to build their own enterprise-grade orchestration platform. Using Kubernetes as the analog for what it can become, it’s easy to see the benefits it can provide to the entire industry.

To understand why observability and OTel’s approach to it are so critical, let’s take a deeper look at telemetry data itself and how it can help organizations transform how they do business.

What is telemetry data?

Telemetry is the process of gathering and transmitting signals (data) emitted by instrumentation code within a system’s components. Traces, metrics, and logs make up most of all telemetry data.

  • Traces result from following a process (for example, an API request or other system activity) from start to finish, showing how services connect. Keeping watch over this pathway is critical to understanding how your ecosystem works, if it’s working effectively, and if any troubleshooting is necessary. Traces consist of individual operations called spans, which include unique identifiers, such as operation name, timestamp, context, attributes, events, and status.
  • Metrics are numerical data points, either counts or measures, that systems can calculate or aggregate over time. Metrics originate from several sources, including infrastructure, hosts, and third-party sources. While logs may not always be accessible, most metrics are readily available via query. Timestamps, values, and even event names can preemptively uncover a growing problem that needs remediation.
  • Logs are important because you’ll naturally want an event-based record of notable anomalies across the system. Structured, unstructured, or plain text, these readable files can tell you the results of any transaction involving an endpoint within your multicloud environment. However, not all logs are inherently reviewable—a problem that’s given rise to external log analysis tools.

Telemetry data becomes observability data when, as noted above, you can “ask meaningful questions, get useful answers, and act effectively on that information.” Making sense of it all requires an observability backend.

How does OpenTelemetry work?

OTel consists of a few components as depicted in the following figure. Let’s take a high-level look at each one from left to right:

OpenTelemetry Components
OpenTelemetry Components (Source: Based on OpenTelemetry: beyond getting started)
  • Specification. Defines a standard telemetry data format and describes how to build instrumentation. This ensures that users have a similar experience, regardless of what language they’re using.
  • Data model. Defines fields for each signal and how they interact. Signals include traces, logs, and metrics.
  • API. Defines the methods used to instrument applications and serves as the entry point for instrumentation. Each language supported by OpenTelemetry has its own API implementation.
  • SDK. Implements the API and also determines how systems generate and correlate their telemetry. Each language supported by OpenTelemetry has its own SDK. Both the APIs and the SDKs are defined in the specification to ensure a consistent experience across implementations.
  • Collector. A vendor-neutral binary used to ingest, transform, and export data to one or more observability backends.
  • OpenTelemetry Protocol (OTLP). A vendor—and tool-agnostic specification for encoding—transmitting and delivering OpenTelemetry data. Telemetry data emitted by the SDK uses OTLP, and many observability backends now support ingesting data in the OTLP format. For those who do not, there are exporters available that convert data from OTLP to a tool-specific format. OTLP supports both HTTP and gRPC.
  • Observability backend. A system or tool where telemetry data collected by OpenTelemetry is sent, stored, and analyzed. It enables organizations to derive meaningful insights and make sense of telemetry data in a cohesive way.

Flexible API/SDK integration

You can decouple the API from the telemetry-generating code with minimal implementation. This decoupling allows your app or library to run with just the API package, without sending telemetry data to the backend. This setup acts as a placeholder until you’re ready to integrate an SDK. When ready, you can choose an SDK that best fits your needs, whether it’s the OpenTelemetry SDK, a vendor-specific one, or a custom-built SDK. This flexibility ensures you can add functionality without significant code changes.

What’s next for OpenTelemetry?

OpenTelemetry is maturing and is fast approaching its graduation as a CNCF project. Traces, logs and most parts of metrics are now considered generally available. The OpenTelemetry project’s goals extend well beyond its current offerings. Exciting initiatives are paving the way for even broader use cases, such as improving digital experiences and enabling detailed insights into application performance through code-level profiling. Let’s check out some highlights of OpenTelemetry’s exciting initiatives, all designed to take observability to the next level.

  • Digital Experience Monitoring. Developers and product teams will soon be able to gather telemetry data directly from user-facing applications. This allows organizations to identify where users face lags or issues, enhancing overall app performance.
  • Code-Level Profiling. OpenTelemetry is also evolving to profile application code in real-time. This provides deeper insights into how specific sections of code behave in production, helping engineers optimize critical parts of their applications.
  • AI Agents. Recently, there’s been an explosion in the need for monitoring AI systems. OpenLLMetry is donating their code to the OpenTelemetry project. If accepted, it will soon become an extension of the OpenTelemetry ecosystem.

How can I contribute to the OTel community?

If you’ve been curious about contributing to OpenTelemetry (OTel) but are unsure where to begin, there are plenty of ways to get involved. Whether you’re a newcomer or a seasoned practitioner, the OpenTelemetry community offers a variety of opportunities suited to different interests and skill levels.

Some ways you can contribute to the OpenTelemetry Project include the OpenTelemetry documentation, OpenTelemetry blog, End User SIG, OpenTelemetry Demo, or a language or component-specific Special Interest Group SIG). No matter how big or small your contributions, they make a difference and are deeply valued.

OpenTelemetry veteran, Adriana Villela, wrote a great article to help you get started.

How does Dynatrace contribute to the Otel community?

Dynatrace is an active member of the OpenTelemetry community. Dynatracers hold key leadership roles as project maintainers or approvers in the following groups:

  • OpenTelemetry Technical Committee
  • OpenTelemetry Specification (Metrics, Semantic Conventions)
  • OpenTelemetry for JavaScript
  • OpenTelemetry Collector
  • OpenTelemetry Demo project
  • OpenTelemetry End User SIG

In fact, Dynatrace has a team dedicated to contributing to OpenTelemetry and ensuring that the Dynatrace platform integrates smoothly with OpenTelemetry data.

Dynatrace and OpenTelemetry together can deliver more value

OpenTelemetry is a key enabler in the observability space, providing a unified framework for collecting telemetry data. However, to unlock the full potential of this data and turn it into actionable insights, a powerful platform to manage, analyze, and visualize it effectively is essential.

Dynatrace is purpose-built to enhance OpenTelemetry’s capabilities. Data plus context are critical to supercharging observability, and with Dynatrace, you’re not just collecting data; you’re gaining a deep understanding of how your systems work and how to optimize them. With seamless integration, advanced analysis across telemetry data, insights into business outcomes, and predictive analytics spanning your entire stack, Dynatrace turns your OpenTelemetry data into actionable intelligence to optimize your systems.

Explore how Dynatrace can transform your OpenTelemetry data into a powerful driver of innovation and business success. Learn more with this video series on getting started with Dynatrace and OpenTelemetry.


Dynatrace Can Do THAT with OpenTelemetry? video thumbnail

Want to explore on your own? Check out the Dynatrace playground.

Want to try Dynatrace with your OpenTelemetry data? Check out our free trial and walk through this Astronomy Shop demo to populate your own data.

The post What is OpenTelemetry?  An open-source standard for logs, metrics, and traces appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/what-is-opentelemetry/feed/ 0
KubeCon EU 2025 key takeaways: Spotlight on cloud-native observability https://www.dynatrace.com/news/blog/kubecon-eu-2025-key-takeaways/ https://www.dynatrace.com/news/blog/kubecon-eu-2025-key-takeaways/#respond Fri, 16 May 2025 01:27:57 +0000 https://www.dynatrace.com/news/?p=69149 KubeCon EU 2025

In April, KubeCon + CloudNativeCon kicked off its 2025 world tour in London. With more than 13,000 attendees, KubeCon EU reaffirmed its reputation as the go-to forum for cloud-native developers and platform engineers. This year’s milestone included the celebration of the CNCF’s 10th birthday and the 11th anniversary of Kubernetes, underscoring the profound impact of […]

The post KubeCon EU 2025 key takeaways: Spotlight on cloud-native observability appeared first on Dynatrace news.

]]>
KubeCon EU 2025

In April, KubeCon + CloudNativeCon kicked off its 2025 world tour in London. With more than 13,000 attendees, KubeCon EU reaffirmed its reputation as the go-to forum for cloud-native developers and platform engineers. This year’s milestone included the celebration of the CNCF’s 10th birthday and the 11th anniversary of Kubernetes, underscoring the profound impact of these technologies on the global tech landscape.

From insights into observability and platform engineering to groundbreaking advancements in security and AI/LLM integrations, the conference offered a wealth of knowledge. For Dynatrace, it was an exciting opportunity to showcase our innovations, collaborate with the open source community, and contribute to discussions shaping the future of cloud-native computing.

Here’s a breakdown of the biggest highlights, emerging trends, and key Dynatrace sessions from KubeCon EU 2025.

Highlights from KubeCon EU 2025

For observability watchers and cloud-native fans, KubeCon EU didn’t disappoint.

1. The unstoppable growth of cloud native

With 9.2 million cloud-native developers worldwide and events expanding from India to Japan and North America, cloud-native technology is scaling faster than ever. The CNCF annual report reveals how essential cloud-native applications have become to industries everywhere.

2. Observability takes center stage

Observability dominated the keynote conversations this year. It’s no longer a “nice-to-have”—it’s a must-have for cloud-native environments. The industry recognizes the need for more basic education around observability, as it becomes crucial for identifying, monitoring, and resolving issues.

3. Kubernetes meets AI and large language models (LLMs)

Kubernetes has solidified its role as the platform of choice for deploying AI models and LLMs. These models demand efficient resource utilization and continuous monitoring to deliver accurate results.

To learn more about how to observe AI and LLMs using Dynatrace, check out our recent Observability Lab, AI and LLM Observability.

4. Platform engineering simplifies Kubernetes complexity

A key trend this year was the rise of Platform Engineering as the ultimate solution to tame Kubernetes complexity. By abstracting away the complex parts of managing K8s, platform engineering enables quicker and more reliable adoption throughout the organization .

5. Security and compliance for cloud-native apps

The EU’s Cyber Resiliency Act brought renewed attention to the importance of security and compliance. Across sessions and keynotes, there was a consensus that these concerns are now everyone’s responsibility, not just security teams.

6. Making observability explainable

One of the most inspiring takeaways came from eBay’s keynote on simplifying observability. They showcased how leveraging LLMs can make complex observability data more understandable, a philosophy that resonates deeply with how Dynatrace’s Davis CoPilot™ simplifies explanations for our users.

Event livestreams

Dynatrace organized two live streams, both pre-and post-KubeCon with cloud native experts to get all the insights about the conference.

Henrik Rexed, Principal Product Manager at Dynatrace, hosted three renowned cloud-native thought leaders on his Is It Observable? Channel, discussing KubeCon EU London and sharing the latest updates. Watch and listen to what Lin Sun, Abdel Sghiouar, and Mauricio Salatino had to say here on YouTube: Observable Live Coding: KubeCon Kickoff.

Make sure to catch the livestream recording of our Dynatrace open source experts Adriana Villela, Alexandra Oberaigner, and Henrik Rexed sharing their insights and takeaways in the Ask Me Anything session on the Dynatrace YouTube channel.

Dynatrace at KubeCon EU 2025

Dynatrace plays a great part in the exciting developments of cloud native. At KubeCon London, 11 Dynatrace innovators presented sessions and posters covering cloud-native technologies and how Dynatrace is helping to shape the future. Here’s an overview of our contributions to the event.

KubeCon EU 2025 Dynatrace speakers list showing 13 separate talks scheduled across three days
List of Dynatracers speaking at KubeCon EU 2025

OpenTelemetry

Putting the Experience in UX: The Importance of Making Data Accessible

Adriana Villela (Dynatrace) and Marino Wijay (Kong Inc.) discuss the importance of making data accessible to all organizational roles, not just IT-focused teams. Using real-life examples, they highlight how a positive experience with data can spark curiosity, unlock new insights, and drive innovation across departments like sales, marketing, finance, and HR.

Customize Your Own OpenTelemetry Collector

Evan Bradley (Dynatrace) and Pablo Baeyens (DataDog) show you how to customize your own OpenTelemetry Collector with the OpenTelemetry Collector Builder (OCB). This tool, developed by the Collector maintainers, lets you include the exact components you need or create your own for unique use cases. The session explores OCB basics, building release pipelines, publishing Docker images, applying hotfixes, and using custom components—plus real-world examples of OCB in action.

How Green Is My OpenTelemetry Collector?

Adriana Villela and Nancy Chauan explore the environmental impact of telemetry and how it contributes to the tech carbon footprint. They introduce the Kepler project, which tracks power consumption metrics in Kubernetes clusters. The session covers what Kepler is, how to deploy it, and a demo on optimizing OpenTelemetry Collector power usage, offering actionable insights for reducing energy consumption and costs.

Smooth Scaling with the OpAMP Supervisor

Evan Bradley (Dynatrace) and Andy Keller (observIQ) show how the OpAMP protocol simplifies managing OpenTelemetry Collectors with seamless remote configuration and control. This session, led by experts from Dynatrace and observIQ, will cover integrating OpAMP into Collector distributions using the new OpAMP Extension and Supervisor. Gain insights into its architecture, features, and see a live demo showcasing centralized configuration, monitoring, and updates.

OTel Sucks (But Also Rocks!)

This talk by Juraci Paixão Kröhling (OllyGarden) and Daniel Dyla (Dynatrace) dives into OpenTelemetry’s challenges and successes, covering SDK configuration improvements, collector performance issues, and evolving semantic conventions. With real-world insights, it offers an honest yet hopeful look at OTel’s growth—perfect for anyone navigating its complexities or appreciating its strengths.

OTel Me How To Get My Open Source Community Taken Seriously

Learn how to build and grow an open source community with Reese Lee (New Relic) and Adriana Villela (Dynatrace). In this session, they’ll share insights from their work with the OpenTelemetry (OTel) community, including strategies for collaboration, driving contributions, and demonstrating business value. Attendees will also hear about key lessons and practical tips to strengthen their own open source projects.

OpenFeature

OpenFeature’s Positive Impact on Confidence at Dynatrace

Simon Schrottner and Todd Baert (Dynatrace) dive into how Dynatrace adopted OpenFeature to address challenges with their fragmented feature flag system, such as unclear use cases and legacy flags. By integrating OpenFeature with OpenTelemetry, they improved feature flag observability, providing Site Reliability Engineers with actionable insights to assess potential impacts confidently. This shift is enhancing their workflows and benefiting the broader developer community.

OpenFeature Updates from the Maintainers

Learn about the latest updates in OpenFeature as presented by Thomas Poignant (Adevinta), Lukas Reining (codecentric AG) and Alexandra Oberaigner (Dynatrace). This session covers new developments like code generation, event tracking, OTEL semantic conventions, distributed flag evaluation, and future plans. Join the discussion and get your questions answered!

Type-safe Feature Flagging in OpenFeature

Learn how to improve feature flagging with OpenFeature, a vendor-agnostic API, in this talk by Michael Beemer (Dynatrace) and Florin-Mihai Anghel (Google). They address common challenges like typos and stale flags in traditional SDKs and demonstrate how the OpenFeature CLI creates type-safe accessors to enhance reliability and the developer experience.

Platform engineering

DORA Metrics in Practice

Learn how to leverage DORA metrics like deployment frequency, lead time for changes, mean time to recovery, and change failure rate to enhance engineering performance. This talk by Danielle Cook, CNCF Ambassador, and Andreas Grabner (Dynatrace) shares insights from Dynatrace’s Internal Development Platform, Juno, covering observability best practices, standard interfaces, and actionable strategies to measure and improve your Software Delivery Lifecycle. Perfect for those curious about or already using DORA metrics!

Cloud-native security

Catch More Hackers with Koney

Poster of Catch Hackers with Koney poster showing the talk abstract, architecture diagram, and deception policies code samples for KubeCon EU

This poster session by Mario Kahlhofer (Dynatrace) and Matteo Golinelli (PhD Student) showed how Koney, a Kubernetes operator, enhances cloud-native app security with honeytokens and deception policies. It automates trap setup, rotation, and teardown, using eBPF to detect and alert on unauthorized access.

To observability and beyond

KubeCon EU 2025 demonstrated that technologies like AI, LLMs, and Kubernetes are reshaping how we think about observability, security, and operational efficiency. At Dynatrace, we’re thrilled to be a part of this transformation, offering AI-powered observability and security that enables teams to innovate faster while staying secure and compliant.

For more great Dynatrace highlights from KubeCon EU 2025, see KubeCon EU 2025 retrospective: Reflections from my sixth KubeCon.

Watch for us at future KubeCon + CloudNativeCon events across the globe in 2025 and beyond!

To learn more about how Dynatrace delivers observability for Kubernetes and cloud-native technologies, read the Kubernetes platform observability best practices guide.

The post KubeCon EU 2025 key takeaways: Spotlight on cloud-native observability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/kubecon-eu-2025-key-takeaways/feed/ 0
Running OpenTelemetry demo app Astronomy Shop with Dynatrace https://www.dynatrace.com/news/blog/opentelemetry-demo-application-with-dynatrace/ https://www.dynatrace.com/news/blog/opentelemetry-demo-application-with-dynatrace/#respond Thu, 13 Mar 2025 16:05:28 +0000 https://www.dynatrace.com/news/?p=68281 OpenTelemetry demo app Astronomy Shop

Many companies are interested in experimenting with the OpenTelemetry observability standard, but have suffered from the lack of a good reference implementation. Now with Astronomy Shop, the OpenTelemetry demo app, organizations can kick the tires using multiple sample microservices that showcase OpenTelemetry features and capabilities. Learn how to deploy and explore OpenTelemetry features using Dynatrace as your full-context OpenTelemetry backend.

The post Running OpenTelemetry demo app Astronomy Shop with Dynatrace appeared first on Dynatrace news.

]]>
OpenTelemetry demo app Astronomy Shop

OpenTelemetry Astronomy Shop is a demo application created by the OpenTelemetry community to showcase the features and capabilities of the popular open-source OpenTelemetry observability standard.

OpenTelemetry provides a common set of tools, APIs, and SDKs to help collect observability signals from applications and infrastructure endpoints. Many companies are interested in experimenting with OpenTelemetry but have struggled with the lack of a good reference implementation.

The Astronomy Shop demo application, which has been actively developed since 2022, solves this problem, and Dynatrace is one of its leading contributors.

OpenTelemetry images: Boy looking through a telescope; tiles with telescope and binoculars images

The OpenTelemetry demo application is a cloud-native e-commerce application made up of multiple microservices. Because it includes examples of 10 programming languages that OpenTelemetry supports with SDKs, the application makes a good reference for developers on how to use OpenTelemetry. You can also use it to test different OpenTelemetry features and evaluate how they appear on backends. Moreover, you can use it as a framework for further customization.

Astronomy Shop reference architecture
Figure 1. OpenTelemetry Astronomy Shop demo application architecture diagram. Courtesy of the OpenTelemetry authors.

To run the demo, you’ll need either a Docker or Kubernetes environment. By default, the demo comes with Jaeger and Prometheus backends, but you can easily configure alternative backends. In this example, we’ll use Dynatrace.

Setting up the OpenTelemetry demo application with Dynatrace as backend

You can ingest OpenTelemetry data in two ways:

Although both methods ingest data, Dynatrace OneAgent helps users automatically discover additional insights about their infrastructure, applications, processes, services, and databases. In this example, we’ll deploy the OpenTelemetry demo application to send telemetry directly to Dynatrace using OTLP so you can see how Dynatrace presents the OTel data without the additional context OneAgent provides.

To set up the demo using Docker, follow the steps below. You can find additional deployment options in the OpenTelemetry demo documentation.

Download and configure Astronomy Shop

  1. First, get and run the Astronomy Shop app from its GitHub repository.
    git clone https://github.com/open-telemetry/opentelemetry-demo.git 
    cd opentelemetry-demo/
  1. Next, configure the demo application to send telemetry to Dynatrace. At the cloned repository’s root, create a new file called  docker-compose.override.yml and paste the following into the file:
    services: 
      otel-collector: 
        environment: 
          - DT_ENDPOINT 
          - DT_API_TOKEN

    This ensures all necessary environment variables that we’ll create in Dynatrace are passed to all services, specifically:

    • DT_ENDPOINT is the OTLP endpoint that is used by the collector to export traces to Dynatrace.
    • DT_API_TOKEN is the token for your Dynatrace environment.
  2. Next, go to src/otel-collector/otelcol-config-extras.yml. This file contains a configuration that will be merged with the Demo app Collector’s default configurator. Copy the following into the file:
    exporters: 
      # otlp/http exporter to Dynatrace.  
      otlphttp/dynatrace:  
        endpoint: "${DT_ENDPOINT}"  
        headers:  
          Authorization: "Api-Token ${DT_API_TOKEN}"  
     
    processors:  
      cumulativetodelta: 
      batch: 
      
    service:  
      pipelines:  
        traces/dynatrace:  
          receivers: [otlp]  
          processors: [batch]  
          exporters: [otlphttp/dynatrace]  
        metrics/dynatrace:  
          receivers: [otlp, spanmetrics]  
          processors: [batch, cumulativetodelta]  
          exporters: [otlphttp/dynatrace]  
        logs/dynatrace:  
          receivers: [otlp]  
          processors: [batch]  
          exporters: [otlphttp/dynatrace]

This configuration will export traces, metrics, and logs to Dynatrace using OTLP. The configuration also includes an optional span metrics connector, which generates Request, Error, and Duration (R.E.D.) metrics from span data. All the needed components are available out of the box in the OpenTelemetry collector contrib distribution, which is included in the demo application.

Set up Dynatrace as backend

In these steps, you’ll set up your Dynatrace account and environment variables.

  1. Create a Dynatrace account. If you don’t have one, you can use a trial account.
  2. Next, create an access token that includes scopes for the following. For details, see Dynatrace API – Tokens and authentication in the Dynatrace documentation.
    • Ingest OpenTelemetry traces (openTelemetryTrace.ingest)
    • Ingest metrics (metrics.ingest)
    • Ingest logs (logs.ingest)
  3. Export the environment variables. Make sure to replace the placeholder values with your token and environment ID. This example illustrates how to pass the token most easily using the terminal. In a production environment, you would configure the token using Kubernetes secrets or other secure secret handling mechanisms.
    export DT_ENDPOINT=https://{your-env-id}.live.dynatrace.com/api/v2/otlp 
    export DT_API_TOKEN=dt0c01.MY_SECRET_TOKEN 
    export OTEL_EXPORTER_OTLP_METRICS_TEMPORALITY_PREFERENCE=delta

Deploy the Astronomy Shop demo application

  1. When everything is prepared, deploy your demo application.
    docker compose up --no-build
    • If you use ARM architecture (for example, a MacBook with Apple silicon), remove the --no-build option to build the images locally.
    • When running the demo for the first time, it takes a couple of minutes to download all the needed images. Afterward, the demo starts instantly, and the load generator automatically begins creating data.
  2. Open the demo application UI to see the app in action:
    http://localhost:8080/

Use Dynatrace to observe traces, metrics, logs, and services

Open your Dynatrace environment in the browser (using the link from creating your Dynatrace account). Once the application is running, you’ll see traces, metrics, and logs start to appear. Dynatrace also automatically detects the services shown in the new Services app.

Next, let’s explore the telemetry data using different Dynatrace apps.

View traces

  1. Using the navigation on the left, open the Distributed Tracing app. In this view, you can analyze traces using the new Distributed Tracing experience. For example, to get traces for the checkout service, you can enter checkout in the Filter requests field, as shown in the following image.
    Dynatrace Distributed Traces app showing checkout traces from the OpenTelemetry demo app
    Figure 2. Analyze checkout traces in the Distributed Tracing app.
  2. If you select one of the traces, you can see the different spans going through multiple services, as shown in the following image.
    Exploring spans of a particular trace from the Astronomy Shop demo in the Dynatrace Distributed Tracing app
    Figure 3. Explore spans of a given trace through various services.

View metrics

  1. Next, take a look at the received metrics by opening the Notebooks app in the left navigation. Select + then select Metrics from the drop-down. Filter metrics using traces.span.metrics.calls to see the metric created by the OTel Collector’s Span Metrics connector.
    Metrics from the OpenTelemetry span metrics connector shown in a line graph
    Figure 4. View metrics created by the OTel Span Metrics connector.
  2. The Span Metrics connector counts the number of spans that each demo app service creates. Use the Split by field with the value service.name to separate the number of spans per service. You can see it’s the front-end proxy that creates the most spans.

View logs

  1. To access logs in Dynatrace, navigate to the Logs app and select Run query.
  2. Next, select one of the log lines to view the available attributes. Select Open trace to connect your log lines with traces and see which request (distributed trace) led to this log line. The magic of the log-trace correlation is the trace_id and span_id attributes that OpenTelemetry instrumentation libraries attach to each log message in an active span.
    Bar chart of the log-trace correlation in an active span from the Astronomy Shop demo app in Dynatrace
    Figure 5. See details of log-trace correlation in an active span

View services

  1. The Services app gives you a high-level view of all your services. In the main view, you can compare the performance and health of each service and detect possible issues.
    List of Astronomy Shop demo services with health line graphs for each in Dynatrace Services app
    Figure 6. Compare the performance and health of each service in the Services app.
  2. When you select one of the services from the list, you get a nice overview of the service and can access its traces, metrics, and logs from a single view.

Using feature flags: Simulating application failures use case

The OpenTelemetry Astronomy Shop demo application contains several built-in use cases that simulate application failures, which you can enable using a feature flag service.

  1. To try one out, go to the feature flag UI:
    http://localhost:8080/feature
  2. Find the feature flag called productCatalogFailure. Turn the feature flag On and select save.
    Feature flag tiles from the OpenTelemetry Astronomy Shop demo application
    Figure 7. Investigate simulated errors using built-in feature flags.
  3. When the product-catalog requests start failing, go to the Distributed Tracing app and select the product-catalog service. You should see some of the traces with Failure status (in red).
  4. Open one of these traces to see more details. You can see that the GetProduct requests are the root cause of these failures. If you select one of the GetProduct spans, you can see the detailed span event showing the reason.
    Dynatrace Distributed Tracing app screen showing the root causes of a failed GetProduct request from the OpenTelemetry demo app
    Figure 8. Investigate the root cause of a failed GetProduct request.

What’s next?

The OpenTelemetry community will continue to improve the demo application so it reflects the latest OpenTelemetry capabilities. Traces, metrics, and logs are already well covered, but interesting enhancements are being made frequently, so stay tuned.

As a top contributor to the OpenTelemetry project since 2020, Dynatrace continues to work with the community and other vendors to enrich the Astronomy Shop project’s capabilities.

The post Running OpenTelemetry demo app Astronomy Shop with Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/opentelemetry-demo-application-with-dynatrace/feed/ 0
Observability as Code: DIY with Crossplane https://www.dynatrace.com/news/blog/observability-as-code-diy-with-crossplane/ https://www.dynatrace.com/news/blog/observability-as-code-diy-with-crossplane/#respond Fri, 14 Feb 2025 15:16:49 +0000 https://www.dynatrace.com/news/?p=67882 Observability as code graphic

In today’s rapidly evolving cloud-native landscape, managing observability efficiently can be a game-changer for any platform. During our recent tech talk at #KCDAustria, Observability as Code – DIY with Crossplane, we demonstrated how the Upjet project can streamline the creation of a custom Crossplane Dynatrace provider, unlocking robust observability capabilities in Kubernetes environments. This blog […]

The post Observability as Code: DIY with Crossplane appeared first on Dynatrace news.

]]>
Observability as code graphic

In today’s rapidly evolving cloud-native landscape, managing observability efficiently can be a game-changer for any platform. During our recent tech talk at #KCDAustria, Observability as Code – DIY with Crossplane, we demonstrated how the Upjet project can streamline the creation of a custom Crossplane Dynatrace provider, unlocking robust observability capabilities in Kubernetes environments. This blog post covers what we did and how we did it.

Crossplane and Dynatrace

In this blog post, we will deploy a monitoring dashboard with alerts and notifications as one simple Kubernetes resource. To help us achieve our goal, we have to set up a Kubernetes cluster and infrastructure components. But before that, we need to discuss some important concepts and patterns.

Configuration as Code

Managing vast amounts of configurations for organizational setups at scale is a hard problem to solve since they span over many tools and providers and have many different contributors.

However, there is a pattern for remedying many of these problems: treating these configurations as declarative code instead of applying changes manually.

Originally dubbed Infrastructure as Code, the pattern can be generalized and used for anything that provides a proper interface—simply put, Configuration as Code.

While many tools and their respective approaches exist to write and apply such configuration, one of the most interesting recent developments is extending Kubernetes and using its readily available REST API and reconciliation loops.

The operator pattern in Kubernetes

The mechanism that Kubernetes provides for interface extension is called the operator pattern. This is a powerful mechanism for automating the management of complex applications. It extends the Kubernetes API via custom resource definitions (CRDs), enabling new object types to be created. These are constantly watched by custom controllers, so-called operators.

Operators are designed with a “reconciliation loop,” meaning they continuously compare a resource’s real state against its desired state. When a resource deviates, the operator brings it back into alignment. This is the essence of Kubernetes automation and declarative infrastructure.

Crossplane and the operator pattern

Crossplane builds on the operator pattern and extends Kubernetes beyond managing containerized workloads. Through the Kubernetes API, it enables you to define and provision—among other things—cloud infrastructure resources such as databases, compute instances, and networking components.

Crossplane providers implement the operator pattern for external systems (for example, AWS, GCP, Azure). When you install one, it installs its external resources as Kubernetes native custom resource definitions (CRDs). The provider’s controller watches for changes in the desired state of objects—instanced from these CRDs, reconciling them to ensure they match your expectations.

Compositions

In Kubernetes, low-level resources are managed by high-level resources. Crossplane also allows you to build high-level resources using the Composition pattern.

An example would be an Application resource that abstracts away details like database, network, and compute needs. A Crossplane composition enables you to build something like this, giving you control over the new interface (also a Kubernetes CRD) and the implementation (which low-level resources are created and how they are created).

Using the Upjet project

Now that we have an overview of the concepts, let’s look at how we implement our demo. When building a custom Crossplane provider, we can take two approaches: build the provider from scratch or leverage an existing tool. We opted for the latter, by using the Upjet project, which automates the creation of Crossplane providers based on existing Terraform providers. Here’s why:

  1. Speed and simplicity: Writing a provider from scratch requires a deep understanding of the external system API and how Crossplane manages resources. Upjet allows us to generate a provider much faster by transforming Terraform provider schemas into Crossplane CRDs, significantly reducing the development time.
  2. Reuse of Terraform providers: Upjet allows us to tap into the vast ecosystem of Terraform providers. Since there are already many well-established Terraform providers for various cloud platforms and services, using Upjet means we don’t have to reinvent the wheel.

Upjet project diagram with Crossplane and Dynatrace

Generating a new Crossplane provider with Upjet

Creating a new Crossplane provider using Upjet is a streamlined process that allows you to extend Crossplane’s capabilities with minimal setup. Follow the official Upjet documentation to get started.

In our demonstration at the KCD Austria tech talk, we showcased how to build a Dynatrace provider. Below are the detailed steps we followed.

Step 1: Adjust the Makefile

The Makefile needs to reference the official Dynatrace Terraform module. This adjustment allows Upjet to use the correct source when generating the provider.

Here’s an example configuration:


export TERRAFORM_PROVIDER_SOURCE ?= dynatrace-oss/dynatrace
export TERRAFORM_PROVIDER_REPO ?= https://github.com/dynatrace-oss/terraform-provider-dynatrace
export TERRAFORM_PROVIDER_VERSION ?= 1.66.0
export TERRAFORM_PROVIDER_DOWNLOAD_NAME ?= terraform-provider-dynatrace
export TERRAFORM_PROVIDER_DOWNLOAD_URL_PREFIX ?= https://releases.hashicorp.com/$(TERRAFORM_PROVIDER_DOWNLOAD_NAME)/$(TERRAFORM_PROVIDER_VERSION)
export TERRAFORM_NATIVE_PROVIDER_BINARY ?= terraform-provider-dynatrace_v1.66.0
export TERRAFORM_DOCS_PATH ?= docs/resources

These settings specify the source, version, and download paths for the Dynatrace Terraform provider that Upjet will wrap as a Crossplane provider.

Step 2: Configure Provider Resources

Set Up the Provider Config

To configure the connection details, we need to modify internal/clients/dynatrace.go to reference the secret structure expected for the provider. In this case, define tenantURL and apiToken for Dynatrace connectivity:


const (
   tenantURL = "dt_env_url"
   apiToken  = "dt_api_token"
)

Then, reference these credentials in the TerraformSetupBuilder:


// TerraformSetupBuilder builds a Terraform setup function, returning provider configuration.
func TerraformSetupBuilder(version, providerSource, providerVersion string) terraform.SetupFn {
    return func(ctx context.Context, client client.Client, mg resource.Managed) (terraform.Setup, error) {
        ...
        // Set credentials in the provider configuration.
        ps.Configuration = map[string]any{}
        if v, ok := creds[tenantURL]; ok {
            ps.Configuration[tenantURL] = v
        }
        if v, ok := creds[apiToken]; ok {
            ps.Configuration[apiToken] = v
        }
    }
}

Define External Name Configurations

To identify external names for resources, update config/external_name.go by adding mappings for the Dynatrace resources:


// ExternalNameConfigs contains all external name configurations for this provider.
var ExternalNameConfigs = map[string]config.ExternalName{
    "dynatrace_alerting":           config.IdentifierFromProvider,
    "dynatrace_email_notification": config.IdentifierFromProvider,
    "dynatrace_json_dashboard":     config.IdentifierFromProvider,
    "dynatrace_metric_events":      config.IdentifierFromProvider,
}

This setup ensures that each resource is correctly identified using the provider’s unique identifier.

Add Custom Configurations for Resources

For each resource, create a corresponding config subfolder and add a config.go file with a Configure function. This function customizes the resource’s configuration and short group name as needed:

config/alerting/config.go


package alerting
import "github.com/crossplane/upjet/pkg/config"
// Configure customizes individual resources.
func Configure(p *config.Provider) {
  p.AddResourceConfigurator("dynatrace_alerting", func(r *config.Resource) {
    r.ShortGroup = "alerting"
  })
}

Repeat this for other resources, such as Dashboard, Event, and Notification.

config/dashboard/config.go


package dashboard
import "github.com/crossplane/upjet/pkg/config"
// Configure customizes individual resources.
func Configure(p *config.Provider) {
  p.AddResourceConfigurator("dynatrace_json_dashboard", func(r *config.Resource) {
    r.ShortGroup = "dashboard"
  })
}

config/event/config.go


package event
import "github.com/crossplane/upjet/pkg/config"
// Configure customizes individual resources.
func Configure(p *config.Provider) {
  p.AddResourceConfigurator("dynatrace_metric_events", func(r *config.Resource) {
    r.ShortGroup = "event"
  })
}

config/notification/config.go


package notification
import "github.com/crossplane/upjet/pkg/config"
// Configure customizes individual resources.
func Configure(p *config.Provider) {
  p.AddResourceConfigurator("dynatrace_email_notification", func(r *config.Resource) {
    r.ShortGroup = "notification"
  })
}

Register Custom Configurations

To ensure these custom configurations are applied, register each Configure function in config/provider.go:


import (
    "github.com/xoanmi/provider-dynatrace/config/alerting"
    "github.com/xoanmi/provider-dynatrace/config/event"
    "github.com/xoanmi/provider-dynatrace/config/dashboard"
    "github.com/xoanmi/provider-dynatrace/config/notification"
)
for _, configure := range []func(provider *ujconfig.Provider){
    alerting.Configure,
    event.Configure,
    notification.Configure,
    dashboard.Configure,
} {
    configure(pc)
}

This setup allows Upjet to apply each resource’s configuration during the provider generation process, customizing each resource group as specified.

Step 3: Generate the Code

Once all the necessary configurations are in place, you’re ready to generate the provider code by running the following command:

The make generate command will use the settings specified in the previous steps to:

  • Generate the Crossplane provider code based on the Terraform provider configurations.
  • Create the necessary Kubernetes Custom Resource Definitions (CRDs) for each resource, allowing Crossplane to manage them.

Running make generate will produce output similar to the following, showing the installation of required tools and the generation of the provider schema and resource CRDs:


➜ make generate
11:38:21 [ .. ] installing terraform darwin-arm64
…
11:38:22 [ OK ] installing terraform darwin-arm64
11:38:22 [ .. ] generating provider schema for dynatrace-oss/dynatrace 1.66.0
11:38:24 [ OK ] generating provider schema for dynatrace-oss/dynatrace 1.66.0
11:38:26 [ .. ] go generate linux_arm64

Generated 4 resources!
11:38:59 [ OK ] go generate linux_arm64
11:38:59 [ .. ] go mod tidy
11:39:00 [ OK ] go mod tidy

➜ tree package/crds
package/crds
├── alerting.crossplane.io_alertings.yaml
├── dashboard.crossplane.io_dashboards.yaml
├── dynatrace.crossplane.io_providerconfigs.yaml
├── dynatrace.crossplane.io_providerconfigusages.yaml
├── dynatrace.crossplane.io_storeconfigs.yaml
├── event.crossplane.io_events.yaml
└── notification.crossplane.io_notifications.yaml

With the generated code and CRDs in place, you can now deploy the provider and begin managing Dynatrace resources in your Kubernetes environment.

Step 4: Deploy and run

In the setup, we’re going to run the operator locally while applying and connecting to a Kubernetes cluster (where this cluster runs doesn’t matter, as long as it’s reachable).

First, we apply the CRDs make generate has created.


kubectl apply -f package/crds

Once this is done, you can run the operator itself.


make run

The last missing part is adding the credentials the operator needs to connect to the chosen Dynatrace tenant. The necessary token is in the Access management documentation.

Create the namespace and a secret containing your brand-new access token.


apiVersion: v1
kind: Secret
metadata:
  name: example-creds
  namespace: crossplane-system
type: Opaque
stringData:
  credentials: |
    {
     "dt_env_url": "https://my-tenant.com",
     "dt_api_token": "my-secret-token"
    }

kubectl create namespace crossplane-system
kubectl apply -f example-creds.yaml

Finally, the last puzzle piece, the ProviderConfig, can be created:


apiVersion: dynatrace.crossplane.io/v1beta1
kind: ProviderConfig
metadata:
  name: default
  namespace: crossplane-system
spec:
  credentials:
    source: Secret
    secretRef:
      name: example-creds
      namespace: crossplane-system
      key: credentials

All done! The previously generated CRDs are now available and their objects in your Kubernetes cluster will result in entities and changes in your Dynatrace tenant. Try it out with a dashboard resource!


apiVersion: dashboard.crossplane.io/v1alpha1
kind: Dashboard
metadata:
  name: example-dashboard
  namespace: crossplane-system
spec:
  forProvider:
    contents: |
      {
        "dashboardMetadata": {
          "name": "Our small example dashboard",
          "owner": "my@mail.com",
          "preset": true,
          "hasConsistentColors": true
      },
      "tiles": [
        {
          More config…
        }
     }

kubectl apply -f example-dashboard.yaml

FAQ

Several key questions were raised during the talk and in the following discussions. We’ve summarized the main points:

Q: Why use Crossplane when Terraform can do the same thing?

A: It’s not a matter of “should” versus “shouldn’t.” If you’re already invested in Kubernetes, Crossplane allows you to manage cloud resources while leveraging the same tooling you use to deploy, maintain, and monitor your applications. This makes Crossplane highly convenient for teams already embedded in the Kubernetes ecosystem.

Additionally, these tools don’t exclude each other. One is used to build platforms, and the other is a command-line tool. Their potential use cases differ quite a lot.

Q: How is state management handled?

A: With the Upjet approach, you’re essentially bridging two worlds. Kubernetes manages the state of each object through its etcd system. Simultaneously, the Crossplane operator runs Terraform in the background, continuously reconciling the state between the Kubernetes Custom Resource (CR) and the Terraform-managed infrastructure.

Q: Can I use the Upjet approach in production?

A: Yes, you can, but remember that the provider uses Terraform under the hood. This means that during each reconciliation loop, a terraform plan and terraform apply run. Due to the nature of these continuous operations, managing a large number of resources this way could demand significant resources.

While the Upject project is very good at translating the provider, some things need to be added manually. The concept of Kubernetes labels simply doesn’t exist in Terraform. If you want to utilize them, you need to implement them yourself.

Q: How does the mapping between Terraform objects and Kubernetes Custom Resources (CRs) work?

A: The mapping is defined in the `/config` folder when configuring the provider. Here, we specify the relationship between the Terraform object and the corresponding Kubernetes CR. Running the `make generate` command triggers the generation of all necessary code, including the API, client, provider, and CRD (Custom Resource Definition). This allows Kubernetes to manage the Terraform-defined resource seamlessly.

Q: I read about the proposal for Crossplane v2.0. Do you know if that will impact the described provider creation process?

A: Recently, the Crossplane developers created a draft for the next version of Crossplane. Here, they talk in-depth about how they want to change composite resources and their structure. This will impact provider creation since they reconcile aforementioned resources. This section discusses the proposed changes. The developers also plan on keeping things backward compatible. For now, we
have to wait and see what the final implementation looks like.

Get started

Crossplane enables a seamless cloud-native approach for managing any cloud resource by extending the Kubernetes API. By leveraging Kubernetes as a control plane and using Crossplane compositions, you can declaratively define and automate your entire observability stack.

It’s easy to get started, all you need to start is

If you’re interested in diving deeper, you can check out the following resources from our session:

We hope this talk inspired you to explore Crossplane for your infrastructure automation needs and provided valuable insights into building observability solutions using the power of Kubernetes and declarative infrastructure.

The post Observability as Code: DIY with Crossplane appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/observability-as-code-diy-with-crossplane/feed/ 0
Dynatrace loves OpenTelemetry https://www.dynatrace.com/news/blog/dynatrace-loves-opentelemetry/ https://www.dynatrace.com/news/blog/dynatrace-loves-opentelemetry/#respond Fri, 14 Feb 2025 00:23:41 +0000 https://www.dynatrace.com/news/?p=67847 Dynatrace and OpenTelemetry

Psst… We have a secret to share: Dynatrace loves OpenTelemetry! And why shouldn’t we? OpenTelemetry (OTel) is an open source framework for generating, ingesting, transforming, and exporting telemetry data. Backed by most of the industry’s major observability vendors, OpenTelemetry has become one of the CNCF’s most active open source projects, second only to Kubernetes. In […]

The post Dynatrace loves OpenTelemetry appeared first on Dynatrace news.

]]>
Dynatrace and OpenTelemetry

Dynatrace loves OpenTelemetry

Psst… We have a secret to share: Dynatrace loves OpenTelemetry! And why shouldn’t we?

OpenTelemetry (OTel) is an open source framework for generating, ingesting, transforming, and exporting telemetry data. Backed by most of the industry’s major observability vendors, OpenTelemetry has become one of the CNCF’s most active open source projects, second only to Kubernetes. In fact, Dynatrace is proud to be one of the leading contributors of OpenTelemetry since the project’s inception.

How Dynatrace supports OpenTelemetry

Dynatrace is an active member of the OpenTelemetry community. Dynatracers hold key leadership roles as project maintainers or approvers in the following groups:

  • OpenTelemetry Technical Committee
  • OpenTelemetry Specification (Metrics, Semantic Conventions)
  • OpenTelemetry for JavaScript
  • OpenTelemetry Collector
  • OpenTelemetry Demo project
  • OpenTelemetry End User SIG

In fact, Dynatrace has a team dedicated to contributing to OpenTelemetry and ensuring that the Dynatrace platform integrates smoothly with OpenTelemetry data.

Dynatrace and OpenTelemery data ingest

Here are some fun facts about Dynatrace ingestion of OpenTelemetry data:

Send OpenTelemetry data to Dynatrace

To send OpenTelemetry data to Dynatrace, you need:

  1. A Dynatrace account
  2. A Dynatrace OTLP endpoint
  3. A Dynatrace API token
  4. An OpenTelemetry-instrumented application

Learn more about sending OpenTelemetry data to Dynatrace.

Explore OpenTelemetry data with Dynatrace

Dynatrace makes unified observability possible by storing all data in Grail™, a unified and purpose-built data lakehouse optimized for storing and analyzing traces, metrics, logs, and more.

Once your data is ingested into Dynatrace, you can use the platform to ask meaningful questions, get useful answers, and act effectively on what you learn. For example, you can:

Analyze your deployed services with automated health analysis and more in the Services app.
Analyze your deployed services with automated health analysis and more in the Services app in Dynatrace

View your traces in the Distributed Tracing app.
View your traces in the Distributed Tracing app in Dynatrace

View your logs, including logs related to the distributed traces, in the Logs app.
View your logs, including logs related to the distributed traces, in the Logs app in Dynatrace

With Notebooks, you can Chart, analyze, set up alerts, and forecast any of your metrics.
With Notebooks, you can Chart, analyze, set up alerts, and forecast any of your metrics in Dynatrace

What’s next?

If you’re interested in learning first-hand what Dynatrace and OpenTelemetry can do together, dive into the following resources:

  1. Check out the Dynatrace Playground with an existing Dynatrace account, or sign up directly for a free tour of the Dynatrace Playground. Once in the Playground, you can use Dynatrace to explore pre-populated OpenTelemetry data.
  2. Configure the OpenTelemetry Demo to send data to Dynatrace, or instrument your own application.
  3. Check out the first video of our new video series, “Dynatrace Can Do THAT with OpenTelemetry?”

Get started with a free trial and ingest your own OpenTelemetry data today

The post Dynatrace loves OpenTelemetry appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-loves-opentelemetry/feed/ 0
Let’s learn how to send OpenTelemetry data to Dynatrace together! https://www.dynatrace.com/news/blog/send-opentelemetry-data-to-dynatrace/ https://www.dynatrace.com/news/blog/send-opentelemetry-data-to-dynatrace/#respond Tue, 21 Jan 2025 20:38:50 +0000 https://www.dynatrace.com/news/?p=67385 OpenTelemetry trends

This blog post will help new and existing customers get started with Dynatrace support for OpenTelemetry. Learn how to send OpenTelemetry data to Dynatrace from an OTel veteran and Dynatrace newbie.

The post Let’s learn how to send OpenTelemetry data to Dynatrace together! appeared first on Dynatrace news.

]]>
OpenTelemetry trends

One of the things I love most about OpenTelemetry (OTel) is that it’s vendor-neutral, which means you can send the same OpenTelemetry data to different vendors. In fact, most of the major Observability vendors out there not only support ingesting OpenTelemetry data but also actively contribute to the project, including Dynatrace. Check out the 2023 OpenTelemetry Journey Report for more info.

Why does this matter? I used to work at another Observability vendor, and many of the OpenTelemetry examples that I played with and blogged about in the last 2 years or so featured sending OTel data to that backend.

Now that I work at Dynatrace, which, by the way, ingests OTLP natively, I wanted to educate myself on how to send OpenTelemetry data to Dynatrace. A great way to learn is to try to run my go-to examples using Dynatrace as the Observability backend. Luckily for me, since OTel is vendor-neutral, all I had to do was reconfigure my OTel Collector to point to Dynatrace to get my examples to work.

Want to learn how? Let’s do it together!

Note: If you’re evaluating multiple vendors, you can send the same data to different vendors at the same time (à la “vendor bake-off”) to help you determine which vendor best suits your organization’s needs.

Prerequisites for sending data to Dynatrace

To send OpenTelemetry data to Dynatrace, you need two pieces of information:

Dynatrace tenant: Each user (or, more likely, organization) is assigned a tenant. When sending OpenTelemetry data from your application to Dynatrace, you need to specify the Dynatrace OTLP endpoint (used by Dynatrace to receive data), which includes your tenant name.

Access token: The access token allows you to send OTel data to your Dynatrace instance. It also specifies what kind of data you’re allowed to send to Dynatrace. You can find more on Dynatrace access tokens here.

Before we get to any of that, you first need a Dynatrace account. If you already have a Dynatrace account, feel free to skip the following section.

Create a Dynatrace account

If you don’t have a Dynatrace account, you can create a free trial account, which is valid for 15 days.

  1. Go here, and select the Free trial button.

    Free trial signup button at dynatrace.com
    Free trial signup button at dynatrace.com
  2. Enter your email, select the Terms of Use checkbox, and then select Continue.
    Enter your email and accept the Terms of Use
    Enter your email and accept the Terms of Use
  3. Enter the rest of the info and select Start free trial.

    Fill in the rest of the form fields
    Fill in the rest of the form fields

    You will receive an email once your account has been created. You will also see a page that looks like the one below. Select Launch Dynatrace to get started.

    Your Dynatrace tenant is ready!
    Your Dynatrace tenant is ready!

    This takes you to the Dynatrace login page.

    Dynatrace login page
    Dynatrace login page

Your Dynatrace tenant

To find your Dynatrace tenant, log into Dynatrace here, and select the Login button at the top right of the page.

This takes you to the sign-in page. Once you sign in, take note of the URL. It should look something like this:

https://<your-dynatrace-tenant>.apps.dynatrace.com

Take note of the value of <your-dynatrace-tenant>, because we’ll need that later.

Create a Dynatrace access token

After confirming that you’re logged into Dynatrace, press ctrl+k, and then type access token. Next, select Access Tokens from the top of the search results.

Access token search
Access token search

On the Access tokens page, select Generate new token.

Dynatrace Access tokens page
Dynatrace Access tokens page

On the Generate new token page, enter:

  • Token name: be sure to give it a descriptive name
  • Expiration date: this is optional
  • Template: Kubernetes Data Ingest

Even if we’re not necessarily using Kubernetes, the Kubernetes Data Ingest template has the token scopes (permissions) that we need to send OpenTelemetry data to Dynatrace, namely:

  • Ingest logs (ingest)
  • Ingest metrics (ingest)
  • Ingest OpenTelemetry traces (ingest)

Find more information on these and other Dynatrace token scopes here.

Once you’re done, select Generate token at the bottom of the page.

Access token generation page
Access token generation page

The next page shows your access token. Be sure to copy and store it somewhere for safekeeping (for example, a secrets manager such as HashiCorp Vault or your cloud provider’s secrets manager) before selecting Done, because after that, it’s gone forever. If you lose that token information, you should delete the old one (not necessary, but highly recommended), and create a new one.

Access token page showing generated token
Access token page showing generated token

Configure the OTel Collector for Dynatrace

Now that you have your tenant info and access token, you can plug this information into your OpenTelemetry Collector configuration.

Note: There are two ways to send OTel data to an Observability backend: (1) direct from the application, or (2) via the OTel Collector. There’s a time and place for each, and you can check out my blog post on OTel Collector Anti-patterns on the OTel Blog to learn more.

Your OTel Collector config YAML file should look something like this:

receivers:
  otlp:
    protocols:
      grpc:
      http:

processors:
  cumulativetodelta:
  batch:

exporters:
  otlphttp:
    endpoint: "https://${DT_TENANT}.live.dynatrace.com/api/v2/otlp"
    headers:
      Authorization: "Api-Token ${DT_API_TOKEN}"
  debug:
    verbosity: detailed

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [batch]
      exporters: [otlphttp,debug]
    metrics:
      receivers: [otlp]
      processors: [cumulativetodelta,batch]
      exporters: [otlphttp,debug]
    logs:
      receivers: [otlp]
      processors: [batch]
      exporters: [otlphttp,debug]

Dynatrace accepts data in the native OpenTelemetry Protocol (OTLP) format via HTTP (gRPC is not yet supported). You need to specify the Dynatrace OTLP endpoint (used by Dynatrace to receive data), which includes your tenant name, ${DT_TENANT}:

https://${DT_TENANT}.live.dynatrace.com/api/v2/otlp

${DT_TENANT} is the value of your Dynatrace tenant name, which you hopefully jotted down and stored in a secrets manager for safekeeping.

Finally, “Dynatrace requires metrics data to be sent with delta temporality, not cumulative temporality”. This means that you’ll need to include the cumulativetodelta processor in:

  • Your Collector configuration (cumulativetodelta)
  • Your metrics pipeline (pipelines.metrics)

Never store your Dynatrace token and tenant name in plain text. Instead, store them in a secrets manager and pull them from the secrets manager at runtime.

Dynatrace and the OTel Operator

If you’re using the OpenTelemetry Operator to send OpenTelemetry data to Dynatrace, you’ll need to configure your OpenTelemetryCollector resource as follows:

apiVersion: opentelemetry.io/v1beta1
kind: OpenTelemetryCollector
metadata:
  name: otelcol
  namespace: opentelemetry
spec:
  mode: statefulset
  image: ghcr.io/dynatrace/dynatrace-otel-collector/dynatrace-otel-collector:0.7.0
  env:
    - name: DT_API_TOKEN
      valueFrom:
        secretKeyRef:
          key: DT_API_TOKEN
          name: otel-collector-secret
    - name: DT_TENANT
      valueFrom:
        secretKeyRef:
          key: DT_TENANT
          name: otel-collector-secret
  config:
    receivers:
      otlp:
        protocols:
          grpc: {}
          http: {}

    processors:
      cumulativetodelta: {}
      batch: {}

    exporters:
      otlphttp:
        endpoint: "https://${DT_TENANT}.live.dynatrace.com/api/v2/otlp"
        headers:
          Authorization: "Api-Token ${DT_API_TOKEN}"
      debug:
        verbosity: detailed

    service:
      pipelines:
        traces:
          receivers: [otlp]
          processors: [batch]
          exporters: [otlphttp,debug]
        metrics:
          receivers: [otlp]
          processors: [cumulativetodelta,batch]
          exporters: [otlphttp,debug]
        logs:
          receivers: [otlp]
          processors: [batch]
          exporters: [otlphttp,debug]

Notice that the spec.config looks the same as what we defined in the otelcol-config.yaml file we saw earlier. The only added thing here is that DT_API_TOKEN and DT_TENANT are environment variables pulled from a Kubernetes Secret. The secret YAML definition looks like this:

apiVersion: v1
kind: Secret
metadata:
 name: otel-collector-secret
 namespace: 
data:
 DT_API_TOKEN: 
 DT_TENANT: 
type: "Opaque"

Both DT_API_TOKEN and DT_TENANT values must be base64 encoded before being added to the secrets YAML. To base64 encode a value and copy the encoded value to your buffer, use this action for easy copy/paste:

echo <value_to_encode> | base64

Remember that storing secrets in Kubernetes (or storing a secrets YAML in version control, for that matter), is not recommended because base64 does not encrypt your data. You should instead consider using sealed secrets or the Kubernetes Secrets Store CSI Driver + your favorite secrets provider. For more info on these better alternatives check out this article.

The Dynatrace OTel Collector Distribution

Many vendors have their own OTel Collector Distributions. These distributions are curated with Collector components that are specific to that vendor. They can be a combination of vendor-developed custom components and components from Collector Core and Contrib. Using vendor-specific distributions ensures you’re using just the Collector components you need, reducing overall bloat. You can learn more here.

Dynatrace also has its own Collector distribution. It features a set of Collector components for sending Observability data to Dynatrace from various sources. It stays up-to-date with upstream components of the opentelemetry-collector and opentelemetry-collector-contrib repositories.

In addition, the Dynatrace Collector distribution offers the following advantages:

  • It is covered by Dynatrace support
  • Collector components are verified by Dynatrace
  • Security patches are independent of OpenTelemetry Collector releases

Try it out!

Want to try sending data to Dynatrace yourself? Then check out my example repo. I created this repo for a talk on the Target Allocator that I gave at KubeCon in March 2024. It has been updated to include instructions on configuring the OpenTelemetryCollector resource to send OTel data to Dynatrace.

OTel Data in Dynatrace

And if you’re curious to see what OTel data looks like in Dynatrace, here are some screenshots of the web UI.

I’m not going in-depth on how to navigate the Dynatrace UI because there are already some great videos on the Dynatrace YouTube channel. I encourage you to check them out for a more in-depth look.

Dynatrace Distributed Tracing UI
Dynatrace Distributed Tracing UI
Dynatrace Logs UI
Dynatrace Logs UI
Dynatrace Notebooks UI showing a metric called ”some_counter_total”
Dynatrace Notebooks UI showing a metric called ”some_counter_total”

Final thoughts

As someone with experience sending OpenTelemetry data to various backends, I found that getting OpenTelemetry data into Dynatrace was fairly straightforward. My only personal hiccup was in generating the application token, but I got that sorted out, and now I’ve passed on my knowledge and highly detailed screenshots along to you.

I have to say that it’s always fun to use a product with fresh eyes, a fresh perspective, and a newbie point of view. There’s nothing quite like it. And, having worked at another observability vendor before, it’s always fun to see the similarities and differences. It’s like learning a new programming language and comparing it to another one that you already know. What a blast!

One final point that I want to make. I don’t want to trivialize things and give you the impression that moving from one observability vendor to another is simply a matter of repointing your OTel Collector from one vendor backend to another. That is only one aspect of a vendor migration, no matter what vendor you’re moving to/from. You also must consider the fact that you’ll likely have a bunch of dashboards, alerts, and whatnot that you created with your original vendor. When you migrate to another vendor, there won’t be a 1:1 translation; so keep that in mind.

But that may be a sacrifice that you’re willing to make because OTel’s vendor neutrality means that all vendors supporting OpenTelemetry are ingesting the same data. What sets them apart is what they do with your data. And if one vendor does something with your data better than another, well, don’t you owe it to yourself to check that out?

What’s Next?

If you’re interested in learning first-hand what Dynatrace and OpenTelemetry can do together, then dive into Dynatrace! You can do so in one of two ways:

  • Check out the Dynatrace Playground to explore Dynatrace using pre-populated OpenTelemetry data
  • Get started with a free trial and ingest your own OpenTelemetry data today!

The post Let’s learn how to send OpenTelemetry data to Dynatrace together! appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/send-opentelemetry-data-to-dynatrace/feed/ 0
When things go sideways: Troubleshooting the OpenTelemetry Operator https://www.dynatrace.com/news/blog/troubleshooting-the-opentelemetry-operator/ https://www.dynatrace.com/news/blog/troubleshooting-the-opentelemetry-operator/#respond Fri, 13 Dec 2024 16:41:22 +0000 https://www.dynatrace.com/news/?p=67050 Kubernetes and OpenTelemetry

Learn the basics of the OpenTelemetry (OTel) Operator and how to troubleshoot when things don’t go according to plan.

The post When things go sideways: Troubleshooting the OpenTelemetry Operator appeared first on Dynatrace news.

]]>
Kubernetes and OpenTelemetry

This blog post was co-written with Reese Lee.

If you already have an application running in Kubernetes and are exploring using OpenTelemetry to gain insights into the health and performance of your app and cluster, you might be interested in an implementation of the Kubernetes Operator called the OpenTelemetry Operator.

As you’ll learn shortly, due to its range of capabilities, the Operator is your go-to for (almost) hassle-free OpenTelemetry management. But, as with any powerful tool, what happens when things go sideways?

In this blog post, you’ll learn about the OpenTelemetry Operator (hereafter referred to as “the Operator”), along with issues commonly encountered across installation, Collector deployment, and auto-instrumentation. You’ll also learn how to resolve these issues, and be better prepared for running the Operator.

Overview of the Operator

Let’s take a closer look at the Operator’s main capabilities.

Managing the Collector

The Operator automates the deployment of your Collector, and makes sure it’s correctly configured and running smoothly within your cluster. The Operator also manages configurations across a fleet of Collectors using Open Agent Management Protocol (OpAMP), which is a network protocol for remotely managing large fleets of data collection agents. Since the protocol is vendor-agnostic, this helps ensure consistent observability settings and simplifies management across agents from different vendors.

Managing Auto-Instrumentation in Pods

The Operator automatically injects and configures auto-instrumentation for your applications, which enables you to collect telemetry data without modifying your source code. If your application isn’t already instrumented with OpenTelemetry, this is a fantastic option to feed two birds with one scone, and start generating and collecting application telemetry.

Installing the Operator

This might seem obvious, but before installing the Operator, you must have a Kubernetes cluster you can install it into, running Kubernetes 1.23+. Check the compatibility matrix for specific version requirements. You can spin up a cluster on your machine using a local Kubernetes tool such as minikube, k0s, or KinD, or use a cluster running on a cloud provider service.

Next, and this is less obvious: You must have a component called cert-manager already installed in that cluster. This piece manages certificates for Kubernetes by making sure the certificates are valid and up to date. You can install both the cert-manager and the Operator via kubectl or a Helm chart.

Note that in either case, you have to wait for cert-manager to finish installing before you install the Operator; otherwise, the operator installation will fail.

Using kubectl

To install cert-manager, run the following command:

kubectl apply -f https://github.com/cert-manager/cert-manager/releases/download/v1.10.0/cert-manager.yaml

Next, install the Operator:

kubectl apply -f https://github.com/open-telemetry/opentelemetry-operator/releases/latest/download/opentelemetry-operator.yaml

Using Helm

To install cert-manager, first add the Helm repository:

helm repo add jetstack https://charts.jetstack.io --force-update

Next, install the cert-manager Helm chart:

helm install \
cert-manager jetstack/cert-manager \
--namespace cert-manager \
--create-namespace \
--version v1.16.1 \
--set crds.enabled=true

Expect the preceding step to take up to a few minutes. You can verify your installation of cert-manager by following the steps in this link, or check the deployment status by running:

kubectl get pods -namespace cert-manager

To install the Operator, note that Helm 3.9+ is required. First, add the repo:

helm repo add open-telemetry https://open-telemetry.github.io/opentelemetry-helm-charts 
helm repo update

Then, install the Operator:

helm install --namespace opentelemetry-operator-system \  
  --create namespace \  
  opentelemetry-operatoropen-telemetry/opentelemetry-operator 

Deploying the OpenTelemetry Collector

Once you have cert-manager and Operator set up in your cluster, you can deploy the Collector. The Collector is a versatile component that’s able to ingest telemetry from a variety of sources, transform the received telemetry in a number of ways based on its configuration, and then export that processed data to any backend that accepts the OpenTelemetry data format (also referred to as OTLP, which stands for OpenTelemetry Protocol).

The Collector can be deployed in several different ways, referred to as “patterns.” Which pattern or patterns you deploy is dependent on your telemetry needs and organizational resources. This topic is out of scope for this blog post, but you can read more about them via this link.

Collector Custom Resource

A custom resource (CR) represents a customization of a specific Kubernetes installation that isn’t necessarily available in a default Kubernetes installation; CRs help make Kubernetes more modular.

The Operator has a CR for managing the deployment of the Collector, called OpenTelemetryCollector. The following is a sample OpenTelemetryCollector resource:

apiVersion: opentelemetry.io/v1beta1
kind: OpenTelemetryCollector
metadata:
  name: otelcol
  namespace: opentelemetry
spec:
  mode: statefulset
  config:
    receivers:
      otlp:
        protocols:
          grpc: {}
          http: {}
      prometheus:
        config:
          scrape_configs:
            - job_name: 'otel-collector'
              scrape_interval: 10s
              static_configs:
              - targets: [ '0.0.0.0:8888' ]

    processors:
      batch: {}

    exporters:
      logging:
        verbosity: detailed

    service:
      pipelines:
        traces:
          receivers: [otlp]
          processors: [batch]
          exporters: [logging]
        metrics:
          receivers: [otlp, prometheus]
          processors: []
          exporters: [logging]
        logs:
          receivers: [otlp]
          processors: [batch]
          exporters: [logging]

There are many configuration options for the OpenTelemetryCollector resource, depending on how you plan on instantiating it; however, the basic configuration requires:

  • mode, which should be one of the following: deployment, sidecar, daemonset, or statefulset. If you leave out mode, it defaults to deployment.
  • config, which may look familiar, because it’s the Collector’s YAML config.

Common Collector deployment issues and troubleshooting tips

If you’re not seeing the data you expect, or you suspect something isn’t working right, try the following troubleshooting tips.

Check that the Collector resources deployed properly

When an OpenTelemetryCollector YAML is deployed, the following objects are created in Kubernetes:

1. OpenTelemetryCollector

2. Collector pod:

  • If you specified non-sidecar mode, look for Deployment, StatefulSet, or DaemonSet resources named <collector_CR_name>-collector-<unique_identifier>).
  • If you specified the mode as sidecar, a Collector sidecar container will be created in an app pod, named otc-container.

3. Target Allocator pod:

  • If you enabled the Target Allocator, look for a resource named <collector_CR_name>-targetallocator-<unique_identifier>.

4. ConfigMap of Collector configurations:

  • If you specified non-sidecar mode, look for Deployment, StatefulSet, or DaemonSet resources named <collector_CR_name>-collector-<unique_identifier>.
  • If you specified the mode as sidecar, note that the Collector config is included as an environment variable.

Thus, when you deploy the OpenTelemetryCollector resource, make sure that the preceding objects are created.

First, confirm that the OpenTelemetryCollector resource was deployed:

kubectl get otelcol -n <namespace>

When you deploy the Collector using the OpenTelemetryCollector resource, it creates a ConfigMap containing the Collector’s configuration YAML. Confirm that the ConfigMap was created in the same namespace as the Collector, and that the configurations themselves are correct.

List your ConfigMaps:

kubectl get configmap -n <namespace> | grep <collector-cr-name>-collector

We also recommend checking your Collector pods by running the appropriate command based on the Collector’s mode:

  • deployment, statefulset, daemonset modes:
kubectl get pods -n <namespace> | grep <collector_cr_name>-collector
  • sidecar mode:
kubectl get pods <pod_name> -n opentelemetry -o jsonpath='{.spec.containers[*].name}'

This will list all the containers created in the pod, including the Collector sidecar container, which includes the Collector config as an environment variable.

Check the Collector CR version

Take a look at the OpenTelemetryCollector CR version you’re using. There are two versions available: v1alpha1:

apiVersion: opentelemetry.io/v1alpha1 
kind: OpenTelemetryCollector 
metadata: 
  name: otelcol 
  namespace: opentelemetry 
spec: 
  mode: statefulset 
  config: | 
    receivers: 
      otlp: 
        protocols: 
          grpc: 
          http: 
 
    processors: 
      batch: 
 
    exporters: 
      otlp: 
        endpoint: "<my_o11y_backend>" 
      logging: 
        verbosity: detailed 
 
    service: 
      pipelines: 
        traces: 
          receivers: [otlp] 
          processors: [batch] 
          exporters: [otlp/ls, logging] 
        metrics: 
          receivers: [otlp, prometheus] 
          processors: 
          exporters: [otlp/ls, logging] 
        logs: 
          receivers: [otlp] 
          processors: [batch] 
          exporters: [otlp/ls, logging] 

and v1beta1:

apiVersion: opentelemetry.io/v1beta1 
kind: OpenTelemetryCollector 
metadata: 
  name: otelcol 
  namespace: opentelemetry 
spec: 
  mode: statefulset 
  config: 
    receivers: 
      otlp: 
        protocols: 
          grpc: {} 
          http: {} 
 
    processors: 
      batch: {} 
 
    exporters: 
      otlp: 
        endpoint: "<my_o11y_backend>" 
      logging: 
        verbosity: detailed 
 
    service: 
      pipelines: 
        traces: 
          receivers: [otlp] 
          processors: [batch] 
          exporters: [otlp/ls, logging] 
        metrics: 
          receivers: [otlp, prometheus] 
          processors: [] 
          exporters: [otlp/ls, logging] 
        logs: 
          receivers: [otlp] 
          processors: [batch] 
          exporters: [otlp/ls, logging] 

There are two main differences between these two API versions:

1. The config sections are different; for v1beta1, the config values are key-value pairs that are part of the CR configuration, whereas for v1alpha1, the config value is one long text string. Keep in mind that the text string still needs to follow YAML formatting.

2. If you’re using v1beta1, you can’t leave the Collector config values empty. You must specify either empty curly braces ({}) for scalar values or empty brackets ([ ]) for arrays. This isn’t necessary if you’re using v1alpha1.

Check the Collector base image

By default, the OpenTelemetryCollector CR uses the core distribution of the Collector. The core distribution is a bare-bones distribution of the Collector for OpenTelemetry developers to develop and test. It contains a base set of components: extensions, connectors, receivers, processors, and exporters.

If you want access to more components than the ones offered by core, you can use the Collector’s Kubernetes distribution instead. This distribution is made specifically to be used in a Kubernetes cluster to monitor Kubernetes and services running in Kubernetes. It contains a subset of components from the core and contrib distributions. Alternatively, you can build your own Collector distribution.

You can set the Collector’s base image by specifying the image attribute in spec.image, as in the following example:

apiVersion: opentelemetry.io/v1beta1 
kind: OpenTelemetryCollector  
metadata: 
  name: otelcol 
  namespace: opentelemetry 
spec: 
  mode: statefulset  
  image: otel/opentelemetry-collector/contrib:0.102.1 
config:  
  receivers:  
    otlp: 
      protocols:  
      grpc: {} 
       http: {} 
  processors:  
    batch: {} 
  exporters:  
    otlp: 
      endpoint: "<olly_backend_endpoint>"

Check your backend vendor’s access requirements

If you’re using a backend vendor to ingest your telemetry data, you’ll likely need to configure an account license key or some kind of access token, which you’ll want to keep confidential.

To store it as a secret and prevent it from appearing as plain text, first create a Kubernetes secret, and Base64-encode it:

apiVersion: v1  
kind: Secret  
metadata: 
  name: otel-collector-secret  
  namespace: opentelemetry 
data: 
  ACCESS_TOKEN: <base64_encoded_token> 
type: "Opaque"

Check your exporter configuration

Confirm that you’ve configured the correct endpoint according to your region in your exporter configuration.

When all else fails…check Kubernetes events

Kubernetes events provide detailed and chronological information about what’s happening within various components of your cluster. To view events for a specific namespace, use:

kubectl get events -n <namespace>

Replace <namespace> with the actual namespace where your OpenTelemetry Operator and resources are deployed.

Instrumentation

Instrumentation is the process of adding code to software to generate telemetry signals–logs, metrics, and traces. You have several options for instrumenting your code with OpenTelemetry, the primary two being code-based and zero-code solutions.

Code-based solutions require you to manually instrument your code using the OpenTelemetry API. While it can take time and effort to implement, this option enables you to gain deep insights and further enhance your telemetry, as you have a high degree of control over what parts of your code are instrumented and how.

To instrument your code without modifying it (or if you’re unable to modify the source code), you can use zero-code solutions (or auto-instrumentation agents). This method uses shims or bytecode agents to intercept your code at runtime or at compile-time to add tracing and metrics instrumentation to the third-party libraries and frameworks you depend on. At the time of publication, auto-instrumentation is currently available for Java, Python, .NET, JavaScript, PHP, and Go. Learn more about zero-code instrumentation at this link.

You can also use both options simultaneously. Some end users opt to start with a zero-code agent and manually insert additional instrumentation, such as adding custom attributes or creating new spans. Alternatively, OpenTelemetry also provides options beyond code-based and zero-code solutions. Learn more at this link.

Zero-code Instrumentation with the Operator

The Operator has a CR called Instrumentation that can automatically inject and configure OpenTelemetry instrumentation into your Kubernetes pods, providing the benefit of zero-code instrumentation for your application. This is currently available for the following: Apache HTTPD, .NET, Go, Java, nginx, Node.js, and Python.

The following is a sample Instrumentation resource definition for a Python service:

apiVersion: opentelemetry.io/v1alpha1  
kind: Instrumentation  
metadata: 
  name: python-instrumentation  
  namespace: application 
spec: 
  env: 
    - name: OTEL_EXPORTER_OTLP_TIMEOUT 
      value: "20" 
    - name: OTEL_TRACES_SAMPLER 
      value: parentbased_traceidratio 
    - name: OTEL_TRACES_SAMPLER_ARG 
      value: "0.85" 
  exporter: 
    endpoint: http://localhost:4317 
  propagators: 
    - tracecontext 
    - baggage  
  sampler: 
    type: parentbased_traceidratio  
    value: "0.25" 
  python:  
    env: 
      - name: OTEL_METRICS_EXPORTER 
        value: otlp_proto_http 
      - name: OTEL_LOGS_EXPORTER 
        value: otlp_proto_http 
      - name: OTEL_PYTHON_LOGGING_AUTO_INSTRUMENTATION_ENABLED 
        value: "true" 
      - name: OTEL_EXPORTER_OTLP_ENDPOINT 
        value: http://localhost: 4318 

You can use a single auto-instrumentation YAML to serve multiple services written in different languages (provided they are supported for auto-instrumentation). List your global environment variables under spec.env, and list language-specific environment variables under spec.<language_name>.env. You can mix and match language-specific environment variable configurations in the same Instrumentation resource.

In order to use the Operator’s auto-instrumentation capability, deploying an Instrumentation resource alone isn’t enough. The auto-instrumentation configuration must be associated with the code being instrumented. This is done by adding an auto-instrumentation annotation in your application’s Deployment YAML, in the template definition section, such as in the following example:

apiVersion: apps/v1 
kind: Deployment  
metadata: 
  name: my-deployment-with-sidecar  
spec: 
  replicas: 1  
  selector: 
    matchLabels: 
      app: my-pod-with-sidecar 
  template: 
    metadata:  
      labels: 
        app: my-pod-with-sidecar 
      annotations: 
        sidecar.opentelemetry.io/inject: "true" 
        instrumentation.opentelemetry.io/inject-python: "true" 
spec: 
  containers: 
    - name: py-otel-server 
      image: otel-python-lab:0.1.0-py-otel-server ports: 
    - containerPort: 8082 
      name: py-server-port 

When the annotation called instrumentation.opentelemetry.io/inject-python is set to true, it tells the Operator to inject Python auto-instrumentation (in this case) into the containers running in this pod. For other languages, simply replace python with the appropriate language name (for example, instrumentation.opentelemetry.io/inject-javafor Java apps). You can disable instrumentation by setting this value to false.

If you have multiple Instrumentation resources, you need to specify which one to use, otherwise the Operator won’t know which one to pick. You can do this as follows:

  • By name. Use this if the Instrumentation resource resides in the same namespaces as the Deployment. For example, opentelemetry.io/inject-java: my-instrumentation will look for an Instrumentation resource called my-instrumentation.
  • By namespace and name. Use this if the Instrumentation resource resides in a different namespace. For example: opentelemetry.io/inject-java: my-namespace/my-instrumentation will look for an Instrumentation resource called my-instrumentation in the namespace my-namespace.

You must deploy the Instrumentation resource before the annotated application; otherwise, your code won’t be automatically instrumented. The Operator injects auto-instrumentation by adding an init container to the application’s pod when it starts up, which means that if the Instrumentation resource isn’t available by the time your service is deployed, the auto-instrumentation will fail.

Common instrumentation issues and troubleshooting tips

If your Collector doesn’t seem to be processing data or if you think the auto-instrumentation isn’t working, try the following steps to troubleshoot and resolve the problem.

Check that the instrumentation resource deployed properly

Run the following command to make sure the Instrumentation resource(s) was created in your Kubernetes cluster:

kubectl describe otelinst -n <namespace>

Confirm the resource deployment order

Double check that your Instrumentation CR is deployed before your Deployment. As we learned earlier, if you’re auto-instrumenting via the Operator, you must deploy the Instrumentation resource before deploying your service’s Deployment resource, because the Deployment will create an init-container for the auto-instrumentation. You should therefore see an auto-instrumentation init-container when you run the following command:

kubectl get pod  -n  \ 
  -o jsonpath='{.spec.initContainers[*].name}' 

Check your auto-instrumentation CR annotations

1- Confirm that there are no typos in the annotations.

2- Confirm that they are in the pod’s metadata definition (spec.template.metadata.annotation), not the deployment’s metadata definition (metadata.annotation), as in the following example:

apiVersion: apps/v1 
kind: Deployment 
metadata: 
  name: py-otel-server 
  namespace: opentelemetry 
  labels: 
    app: my-app 
    app.kubernetes.io/name: py-otel-server 
spec: 
  replicas: 1 
  selector: 
    matchLabels: 
      app: my-app 
      app.kubernetes.io/name: py-otel-server 
  template: 
    metadata: 
      labels: 
        app: my-app 
        app.kubernetes.io/name: py-otel-server 
      annotations: 
        instrumentation.opentelemetry.io/inject-python: "true" 
    spec: 
      containers: 
      - name: py-otel-server 
        image: otel-target-allocator-talk:0.1.0-py-otel-server 
        imagePullPolicy: IfNotPresent 
        ports: 
        - containerPort: 8082 
          name: py-server-port 
        env: 
          - name: OTEL_RESOURCE_ATTRIBUTES 
            value: service.name=py-otel-server,service.version=0.1.0 

Check your endpoint configurations

The endpoint, configured in the following example under spec.exporter.endpoint, refers to the destination for your telemetry within your Kubernetes cluster:

apiVersion: opentelemetry.io/v1alpha1 
kind: Instrumentation 
metadata: 
  name: python-instrumentation 
  namespace: opentelemetry 
spec: 
  exporter: 
    endpoint: http://otelcol-collector.opentelemetry.svc.cluster.local:4318 
  env: 
  propagators: 
    - tracecontext 
    - baggage 
  python: 
    env: 
      - name: OTEL_METRICS_EXPORTER 
        value: console,otlp_proto_http 
      - name: OTEL_LOGS_EXPORTER 
        value: otlp_proto_http 
      - name: OTEL_PYTHON_LOGGING_AUTO_INSTRUMENTATION_ENABLED 
        value: "true" 

The spec.exporter.endpoint configuration in the Instrumentation resource allows you to define the destination for your telemetry data. If you omit it, it defaults to http://localhost:4317.

If you’re sending out your telemetry to a Collector, the value of spec.exporter.endpoint must reference the name of your Collector Service.

Looking at the example above, otel-collector is the name of the OTel Collector Kubernetes Service.

In addition, if the Collector is running in a different namespace, you must append opentelemetry.svc.cluster.localto the Collector’s service name, where opentelemetry is the namespace in which my Collector happens to be deployed to. It can be any namespace of your choosing.

Finally, make sure that you are using the right Collector port. Normally, you can choose either 4317 (gRPC) or 4318(HTTP); however, for Python auto-instrumentation, you can only use 4318. Confirm whether there are similar caveats for the language(s) you’re using.

Note: If you’re deploying your Collector as a Sidecar, your endpoint needs to be  http://localhost:4317 or http://localhost:4318 (remember: it has to be 4318 for Python).

When all else fails…check the Operator logs

Run the following command to check the Operator logs for any occurrences of error in the log messages:

kubectl logs -l app.kubernetes.io/name=opentelemetry-operator \ 
  --container manager \ 
  -n opentelemetry-operator-system --follow 

Note that the above only applies if you have admin access to your Kubernetes cluster. If you don’t, you can still tell what’s going on by checking your Kubernetes event log, just like we did when troubleshooting issues with the OpenTelemetryCollector resource:

kubectl get events -n <namespace>

Summary

The OpenTelemetry Operator manages the deployment and configuration of one or more Collectors, and injects and configures zero-code instrumentation solutions into your Kubernetes pods. This enables you to get started with OpenTelemetry instrumentation, and you can further enhance your telemetry by adding manual instrumentation to your application.

In this blog post, you learned the ins and outs of the Operator, from common installation hurdles to resolving auto-instrumentation and Collector deployment issues. With detailed installation steps and troubleshooting tips, you’re now equipped to leverage the Operator effectively for the deployment, configuration, and management of your Collectors and auto-instrumentation of supported libraries.

This blog post is based on a talk that Adriana and Reese did at KubeCon North America’s 2024 co-located event, Observability Day. You can check out the recording of the talk here:

When Things Go Sideways: Troubleshooting the OTel Operator – Adriana Villela & Reese Lee

The post When things go sideways: Troubleshooting the OpenTelemetry Operator appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/troubleshooting-the-opentelemetry-operator/feed/ 0
OpenTelemetry histograms reveal patterns, outliers, and trends https://www.dynatrace.com/news/blog/opentelemetry-histograms-reveal-patterns-outliers-and-trends/ https://www.dynatrace.com/news/blog/opentelemetry-histograms-reveal-patterns-outliers-and-trends/#respond Thu, 31 Oct 2024 08:00:36 +0000 https://www.dynatrace.com/news/?p=66111 Dynatrace introduces support for OpenTelemetry histograms

Dynatrace introduces support for OpenTelemetry histograms, which visualize and make it easier to understand the distribution of data. These histograms enable, for example, response time analysis for services and help to define and monitor service-level objectives that can be alerted on.

The post OpenTelemetry histograms reveal patterns, outliers, and trends appeared first on Dynatrace news.

]]>
Dynatrace introduces support for OpenTelemetry histograms

Imagine you’re using a lot of OpenTelemetry and Prometheus metrics on a crucial platform. You’re gathering a lot of data, but you can’t make sense of it. You need to visualize the distribution of your measurements to identify patterns, outliers, and trends. But there’s a problem: Your current tools don’t support histograms.

Incorporating histograms is not just a technical upgrade—it’s a necessity for any observability professional. By starting with histograms, you can unlock deeper insights and drive more informed decisions in your projects.

We’re excited to announce that Dynatrace has introduced support for OpenTelemetry histograms in connection with the new visualization options in Dashboards and Notebooks. The histograms are supported starting from Dynatrace version 1.301. OpenTelemetry histograms complement the Distributed Tracing app, which uses histograms as the default visualization tool for response times.

In this blog, we will focus on histograms and why to use them. We will cover their main value and possibilities in OpenTelemetry.

Histograms are commonly used to define and monitor service-level objectives (SLOs). They can help determine the percentage of requests that meet a specific response-time threshold, which is essential for maintaining service quality.

In practice, histograms are useful when the measurement distribution is relevant and the data sets are large. Teams can also change queries to get answers on already-collected data without needing to redefine metrics or wait for new data to accumulate.

OpenTelemetry histograms

Breaking down the benefits of OpenTelemetry histograms

OpenTelemetry instrumentation automatically generates histograms for HTTP client and server request durations. This feature, available by default for OTel-instrumented services, gives users a standard way to consistently measure and compare response times across different services.

Moreover, the OpenTelemetry Collector can measure service span durations, categorized by span names, span kinds, and status codes. The span metrics connector creates these measurements and presents them as histograms, which you can analyze in Dynatrace for deeper insights.

Histograms also enhance the self-monitoring capabilities of the Collector. It reports batch sizes and HTTP/RPC measurements of its own pipelines as histograms, providing valuable metrics for performance monitoring. This self-monitoring aspect is crucial for maintaining the health and efficiency of the Collector itself, ensuring that it can handle the demands of large-scale data collection and processing without degradation.

Additionally, the Collector supports converting Prometheus and StatsD histograms into the OpenTelemetry protocol (OTLP), making them compatible with Dynatrace. By exporting metrics from different sources into a single platform, teams can achieve a holistic view of their system’s performance, facilitating proactive issue resolution and faster decision-making.

Percentiles to simplify analysis

Percentiles are statistical measures that divide a data set into 100 equal parts, providing a way to interpret specific points within your histograms. For instance, the 90th percentile (p90) is the value below which 90% of the data falls.

In practical applications, percentiles are particularly useful for web performance analysis. By examining the p90, you can identify the maximum response time experienced by 90% of users. This insight is crucial for optimizing performance for the majority of users. However, it also highlights that the remaining 10% of users experience longer wait times, which could lead to dissatisfaction.

With the Dynatrace Grail data lakehouse, extracting percentiles from histograms is straightforward, especially when using Notebooks. You can seamlessly integrate percentile graphs into dashboards, providing clear and actionable insights.

OpenTelemetry histograms with Dynatrace Grail

Support for explicit and exponential histograms

The first metrics API/SDK release in the OpenTelemetry project introduced histograms with explicit bucket boundaries. These histograms are very popular and are also widely used by Prometheus. Dynatrace now fully supports them.

Later, OpenTelemetry introduced exponential histograms, with each consecutive bucket exponentially larger than the previous one. These histograms are more efficient in carrying a high dynamic range of different values and ensure that the relative error for every bucket remains stable. Dynatrace now supports exponential histograms by calculating histogram summaries (min, max, sum, count). But for now, percentile calculation and buckets are available only for explicit bucket histograms.

Try OpenTelemetry histograms

To experiment with OpenTelemetry histograms, you can deploy the OpenTelemetry Demo Application (Astronomy shop) with the span metrics connector. See this blog about exporting the data from the demo app to Dynatrace.

To learn more about the histograms in Dynatrace, see Histogram Visualization in Dynatrace docs.

For easy analysis of trace data with histograms, check out the new Distributed Tracing app. You can also check out this demo: Transform OpenTelemetry data into actionable insights.

As a leading contributor to the OpenTelemetry project, Dynatrace is committed to advancing its features and maximizing its value. By collaborating with the community and other vendors, Dynatrace ensures that OpenTelemetry remains cutting-edge, accessible, and user-friendly for everyone.

The post OpenTelemetry histograms reveal patterns, outliers, and trends appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/opentelemetry-histograms-reveal-patterns-outliers-and-trends/feed/ 0
The Force of interfaces and meaningful class names: Don’t Be a Sith https://www.dynatrace.com/news/blog/the-force-of-interfaces-and-meaningful-class-names-dont-be-a-sith/ https://www.dynatrace.com/news/blog/the-force-of-interfaces-and-meaningful-class-names-dont-be-a-sith/#respond Tue, 09 Jul 2024 19:13:34 +0000 https://www.dynatrace.com/news/?p=64658 Class names

In a galaxy far, far away, there’s a disturbance in the Force. It’s a practice that’s as limiting as the Sith Rule of Two: “Always two there are, no more, no less. A master and an apprentice.” The convention I’m referring to is naming Java classes with a trailing Impl. The dark side of IMPL Imagine […]

The post The Force of interfaces and meaningful class names: Don’t Be a Sith appeared first on Dynatrace news.

]]>
Class names

In a galaxy far, far away, there’s a disturbance in the Force. It’s a practice that’s as limiting as the Sith Rule of Two: “Always two there are, no more, no less. A master and an apprentice.” The convention I’m referring to is naming Java classes with a trailing Impl.

The dark side of IMPL

Imagine you have an interface, InputHandler. The dark side tempts you to name the implementation InputHandlerImpl. This is a trap! It’s as restrictive as the Sith philosophy, limiting the potential of your code because it implicitly suggests there should only be one implementation of that interface.

The IMPL suffix lures developers into the wrongful thinking that only one implementation of that interface should exist.
The IMPL suffix lures developers into the wrongful thinking that only one implementation of that interface should exist.

Embrace the light side with descriptive naming

Instead of succumbing to the dark side, let’s embrace the light side of the Force. The name of the implementing class should start with Default. Better yet, it should include some information about why this class exists.

So, instead of InputHandlerImpl, maybe consider DefaultInputHandler. This small difference opens your mind and leaves room for further implementations, if necessary. Even better options are HttpInputHandler or StreamingInputHandler, as these document technical details that add real value to the name.

The power of many can bring about change.
The power of many can bring about change.

This approach aligns more with the Jedi way, allowing multiple Padawans to learn and grow, just as multiple classes can implement the same interface, each with unique characteristics.

This method not only keeps your mind open to additional implementations but also provides more insight into the technical details of the code. Just browsing through the classes of the code allows anyone to learn about the protocols and technologies automatically.

Don’t be a Sith, be a Jedi

Remember, restricting interfaces to one implementing class is like the Sith, where there can only be one master and one apprentice. But in software development, we want to be more like the Jedi, embracing diversity and learning from different implementations.

So, fellow developers, let’s not be Sith, limiting ourselves to a single Impl. Let’s be Jedi, mindfully leveraging the vast possibilities of descriptive class naming. May the Force (and good coding practices) be with you!

Learn more about software engineering at Dynatrace.

The post The Force of interfaces and meaningful class names: Don’t Be a Sith appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/the-force-of-interfaces-and-meaningful-class-names-dont-be-a-sith/feed/ 0
Best practices for ingesting logs, traces, and metrics with Fluent Bit 3.0 https://www.dynatrace.com/news/blog/best-practices-for-fluent-bit-3-0/ https://www.dynatrace.com/news/blog/best-practices-for-fluent-bit-3-0/#respond Tue, 07 May 2024 08:40:36 +0000 https://www.dynatrace.com/news/?p=63944 fluentbit memory usage

This blog post will introduce you to Fluent Bit 3.0 and some best practices for using it in your observability pipeline.

The post Best practices for ingesting logs, traces, and metrics with Fluent Bit 3.0 appeared first on Dynatrace news.

]]>
fluentbit memory usage

The recent release of Fluent Bit 3.0 offers some new opportunities for Fluent Bit best practices. Let’s take a look at Fluent Bit and what’s new for v3.

What is Fluent Bit?

Fluent Bit is a telemetry agent designed to receive data (logs, traces, and metrics), process or modify it, and export it to a destination. Fluent Bit can serve as a proxy before you send data to Dynatrace or similar. However, you can also use Fluent Bit as a processor because you can perform various actions on the data. Fluent Bit was designed to help you adjust your data and add the proper context, which can be helpful in the observability backend.

What’s the difference between Fluent Bit and Fluentd?

Fluent Bit and Fluentd were created for the same purpose: collecting and processing logs, traces, and metrics. However, Fluent Bit was designed to be lightweight, multi-threaded, and run on edge devices.

Fluent Bit was created before Kubernetes existed when the internet of things (IoT) was a new buzzword. Unfortunately, their market prediction wasn’t correct; the cloud became more successful than IOT.

Learn more about configuring Fluent Bit to collect logs for your Kubernetes cluster and collecting logs with Fluentd.

What’s new in Fluent Bit 3.0

Fluent Bit recently released version 3.0, which offers a range of updates:

  • HTTP/2 support: Fluentbit now supports HTTP/2, enabling efficient data transmission with Gzip compression for OpenTelemetry data, enhancing pipeline performance.
  • New processors: Introducing new processors, including Metric Selector and Content Modifier, for selective data processing and metadata adjustment, improving data relevance and storage efficiency.
  • SQL processor: Integration of SQL processing capabilities for precise data selection and routing, allowing users to filter logs and traces based on specific criteria for optimized backend storage.
  • Enhanced filtering: Streamlined data collection with selective metrics inclusion/exclusion, efficient metadata adjustment, and precise data routing, ensuring only relevant data is transmitted to the observability backend.
  • Scalability and efficiency: Improved scalability and performance enhancements, optimized compression, and efficient data routing contribute to a more streamlined and cost-effective observability solution with Fluentbit 3.0.

In the following paragraphs, we will introduce a few best practices, particularly when using Fluent Bit 3.0. If you’d like to learn more about the updates, don’t miss Henrik’s video: Fluent Bit 3.0 Observability: Elevating Logs, Metrics, and Traces!

Fluent Bit best practices for v3.0

Here are some best practices for using Fluent Bit 3.0, as recommended by our internal expert, Henrik Rexed.

Understand what you want to accomplish

Before using Fluent Bit, you should clearly understand what you want to accomplish with your data. If you have logs, Fluent Bit is the right choice. For metrics, Fluent Bit has some limitations. And for traces, sampling won’t be possible. Be sure you can achieve what you want with Fluent Bit; otherwise, you might need to look into other options. Ask yourself, how much data should Fluent Bit process? Would it be better to send the data somewhere else, like Dynatrace, and let Grail process it?

Use the Expect plugin

Start small and test that you’re achieving what you expect by using the Expect plugin, a gatekeeper that helps your pipeline be more efficient by blocking content that doesn’t follow the rules you put in place.

Before reading the expected logs, you can either create a file with log data designed for testing or use the dummy input plugin that sends dummy data to the pipeline (see example below).

pipeline:

 inputs:

   - name: dummy

     dummy: '{"message": "custom dummy"}'

 outputs:

   - name: stdout

     match: '*'

Apply modifications directly at the source

Attaching the processor to the input or output sections allows you to apply modifications directly at the source, making you more efficient. Localized modifications are thus possible before logs go through the parser, which means you can:

  • Make global rules
  • Be more tactical
  • Save CPU time
  • Have a more robust and reliable pipeline

Processing at the source relieves pressure on the Fluent Bit agent.

Other types of modifications

Traces

It might be a rare use case, but if you want to obfuscate some traces, you can use the content_modifier. The content modifier can only be used for traces and logs, not metrics.

Metrics

Use the label processor to remove metrics, add them, and reduce cardinality and metric costs. The only thing missing is converting the metric (cumulative and delta).

Logs

The SQL processor makes your pipeline experience easier since you can use select to filter precisely which logs you want to change.

Understand the concept of tags

Your experience with Fluent Bit will improve if you learn how tags function, especially if you’ve never used Fluent before. Tags make routing possible and are set in the configuration of the input definitions. They can be used to apply specific operations only to specific data.

For example, you’ve ingested OpenTelemetry metrics, logs, and traces, so you might have:

  • Otlp.metrics
  • Otlp.logs

Using the tag “metrics,” you can target just the otlp.metrics for certain operations.

There’s also a filter plugin called rewrite_tag that, as the name suggests, you can reuse to re-emit a record under a new tag.

Use the service section to check your pipeline’s health

In the service section, you can enable two crucial things: the Health_Check and the http_service. These help you get more insight into the health of your pipeline with Kubernetes. See the example below.

apiVersion: v1

kind: ConfigMap

metadata:

 annotations:

   meta.helm.sh/release-name: fluent-bit

   meta.helm.sh/release-namespace: fluentbit

 labels:

   app.kubernetes.io/instance: fluent-bit

   app.kubernetes.io/managed-by: Helm

   app.kubernetes.io/name: fluent-bit

   app.kubernetes.io/version: 2.2.1

   helm.sh/chart: fluent-bit-0.42.0

 name: fluent-bit

 namespace: fluentbit

data:

 custom_parsers.conf: |

   [PARSER]

       Name docker_no_time

       Format json

       Time_Keep Off

       Time_Key time

       Time_Format %Y-%m-%dT%H:%M:%S.%L

 fluent-bit.yaml: |

   service:

     http_server: "on"

     Health_Check: "on"


   pipeline:

     inputs:

       - name: tail

         path: /var/log/containers/*.log

         multiline.parser: docker, cri

         tag: kube.*

         mem_Buf_Limit: 5MB

         skip_Long_Lines: On

         processors:

           logs:

             - name: content_modifier

               action: insert

               key: k8s.cluster.name

               value: ${CLUSTERNAME}

             - name: content_modifier

               action: insert

               key: dt.kubernetes.cluster.id

               value: ${CLUSTER_ID}

             - name: content_modifier

               context: attributes

               action: upsert

               key: "agent"

               value: "fluentbitv3"

      

       - name: fluentbit_metrics

         tag:  metric.fluent

         scrape_interval: 2

   

     filters:


       - name: kubernetes

         match: kube.*

         merge_log: on

         keep_log: off

         k8s-logging.parser : on

         k8S-logging.exclude: on




       - name: nest

         match: kube.*

         operation: lift

         nested_under: kubernetes

         add_prefix :  kubernetes_

       - name: nest

         match: kube.*

         operation: lift

         nested_under: kubernetes_labels

       - name: modify

         match: kube.*

         rename:

           - log content

           - kubernetes_pod_name k8s.pod.name

           - kubernetes_namespace_name k8s.namespace.name

           - kubernetes_container_name k8S.container.name

           - kubernetes_pod_id k8s.pod.uid

         remove:

            - kubernetes_container_image

            - kubernetes_docker_id

            - kubernetes_annotations

            - kubernetes_host

            - time

            - kubernetes_container_hash

      

       - name: throttle

         match: "*"

         rate:     800

         window:   3

         print_Status: true

         interval: 30s

     


     outputs:

      

       - name: opentelemetry

         host: ${DT_ENDPOINT_HOST}

         port: 443

         match: "kube.*"

         metrics_uri: /api/v2/otlp/v1/metrics

         traces_uri:  /api/v2/otlp/v1/traces

         logs_uri: /api/v2/otlp/v1/logs

         log_response_payload: true

         tls:  On

         tls.verify: Off

         header:

           - Authorization Api-Token ${DT_API_TOKEN}

           - Content-type application/x-protobuf

       - name: prometheus_exporter

         match: metric.*

         host: 0.0.0.0

         port: 2021

Enabling the http_server also exposes many API endpoints.

Utilize Hot Reload to save time with changes

By default, any change you make requires an agent restart, which can be time-consuming. If you enable Hot Reload, an HTTP request is sent to reload Fluent Bit and read the pipeline files, which makes the whole process smoother. It’s particularly great if you’re in a bare metal environment.

Expose data for Prometheus with Fluent Bit Metrics

If you want to expose data from a Prometheus standpoint, use the Fluent Bit Metrics plugin, which asks Fluent Bit to become a Prometheus exporter. The plugin is excellent because it produces interesting metrics.

service:


   flush: 1


   log_level: info


pipeline:


   inputs:


       - name: fluentbit_metrics


         tag: internal_metrics


         scrape_interval: 2


   outputs:


       - name: prometheus_exporter


         match: internal_metrics


         host: 0.0.0.0


         port: 2021

Mind your memory and resources

When you modify your data and reject what you don’t need, a retry logic keeps the rejected data in memory. However, if too many exceptions exist, the Fluent Bit agent will become unstable and crash. So it’s important to regularly look at the stdout of Fluent Bit and remove the noise; otherwise, the logs may become unreliable.

In every input plugin for logs and traces, we see a (memory) buffer. Data can be ingested faster than it can be sent to its destination, which can create backpressure, leading to high memory consumption. Plan to avoid out-of-memory situations. By default, you have a storage type memory, but you may exceed this buffer limit if you have a lot of data. It’s crucial to finetune this memory buffer limit to avoid instability.

On a similar note, pay attention to the resource usage of the agent (CPU, behavior, throttling) and the pipeline metrics. Create a dashboard to see those metrics, as you can see in the example below from Dynatrace:

Fluent Bit dashboard in Dynatrace screenshot

An overview like this can help you see if you’re losing data. For example, “I’m receiving 100 metrics, but only 30 are being outputted, which means I am losing 70% of my data.” Is this intentional? Are there a lot of retries? Retries could be the cause of this issue.

In the dashboard, the Fluent Bit logs are very telling. The logs are rich in information and can help you identify issues that might cause Fluent Bit to be unstable. For example, you can see if there is a major difference between the data being sent into Fluent Bit and how much arrives in Dynatrace.

Fluent Bit dashboard in Dynatrace screenshot

Conclusion

In this blog post, we delved into the latest updates in Fluent Bit 3.0 and outlined some best practices for leveraging its capabilities effectively within your observability pipeline.

Learn more about Fluent Bit 3.0 on this YouTube video created by Henrik: Fluent Bit 3.0 Observability: Elevating Logs, Metrics, and Traces!

View Henrik’s YouTube video, “Fluent Bit 3.0 Observability: Elevating Logs, Metrics, and Traces!”

The post Best practices for ingesting logs, traces, and metrics with Fluent Bit 3.0 appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/best-practices-for-fluent-bit-3-0/feed/ 0
Mastering Kubernetes deployments with Keptn: a comprehensive guide to enhanced visibility https://www.dynatrace.com/news/blog/mastering-kubernetes-deployments-with-keptn-a-comprehensive-guide-to-enhanced-visibility/ https://www.dynatrace.com/news/blog/mastering-kubernetes-deployments-with-keptn-a-comprehensive-guide-to-enhanced-visibility/#respond Wed, 20 Mar 2024 15:34:52 +0000 https://www.dynatrace.com/news/?p=63168 Keptn logo

Deploying software in Kubernetes is often viewed as a straightforward process—just use kubectl or a GitOps solution like ArgoCD to deploy a YAML file, and you’re all set, right? Unfortunately, Kubernetes deployments can be fraught with challenges beyond the surface level. Numerous hurdles can hinder successful deployments, from resource constraints to external dependencies and monitoring […]

The post Mastering Kubernetes deployments with Keptn: a comprehensive guide to enhanced visibility appeared first on Dynatrace news.

]]>
Keptn logo

Deploying software in Kubernetes is often viewed as a straightforward process—just use kubectl or a GitOps solution like ArgoCD to deploy a YAML file, and you’re all set, right? Unfortunately, Kubernetes deployments can be fraught with challenges beyond the surface level. Numerous hurdles can hinder successful deployments, from resource constraints to external dependencies and monitoring inadequacies. In this article, we’ll explore these challenges in detail and introduce Keptn, an open source project that addresses these issues, enhancing Kubernetes observability for smoother and more efficient deployments.

Understanding Kubernetes deployment challenges

Deploying software using Kubernetes is, on the surface, easy, but there is a lot more to ensure a healthy deployment. There are a lot of potential problems that can prevent the successful deployment of a Kubernetes application, such as:

Resource constraints

In Kubernetes, managing resources efficiently is crucial to prevent performance bottlenecks and application failures. Insufficient CPU and memory allocation to pods can lead to resource contention and stop Pods from being created.

External dependencies

Many applications rely on external services, such as databases, APIs, or third-party services. Ensuring seamless connectivity to these dependencies during deployment is essential for application stability. Consider a scenario where a web application depends on an external payment gateway. Any connectivity issues with the payment gateway during deployment could result in transaction failures and impact the application’s functionality.

Infrastructure health

The underlying infrastructure’s health directly impacts application availability and performance. Vulnerabilities or hardware failures can disrupt deployments and compromise application security. For instance, if a Kubernetes cluster experiences a hardware failure during deployment, it can lead to service disruptions and affect the user experience.

Monitoring inadequacies

Comprehensive monitoring is essential for detecting and resolving issues during deployment. Inadequate monitoring tools or configurations can result in operational blind spots, delaying issue detection and resolution.

Performance degradation

Deploying new versions of applications can sometimes lead to unexpected performance degradation. Changes in application code or configurations can impact performance metrics, affecting user experience and application functionality. For instance, deploying a new version of a web application that introduces inefficient database queries could lead to increased response times and decreased user satisfaction.

This is the problem domain on which Keptn can help Platform Engineers provide a solution and guard rails for teams to deploy software that works perfectly.

Addressing Kubernetes deployment challenges with Keptn

Using Keptn, combined with your standard deployment tooling or practices,

you can move from “I guess it’s okay” to “I know it’s okay”. Keptn allows you to wrap governance and automated checks around the deployment process to ensure that the end-to-end deployment process is healthy and that your application meets the SLOs you’ve defined. This is done by configuring Keptn to execute KeptnTasks and KeptnEvaluations before and after deploying a Kubernetes Workload.

With these, you can implement checks that ensure your infrastructure is in a healthy state before a new version of a deployment is rolled out. This way, you can avoid the problems mentioned above by:

  • Verifying that your cluster has enough resources to run a workload.
  • Checking the reachability of external services that your workloads might require.
  • Ensuring that your monitoring solution monitors your cluster.
  • Verifying that there are no open problems within your cluster.

After successfully deploying a workload, Keptn also provides the means for post-deployment checks. These can be crucial in ensuring that the performance of your newly deployed workload meets your expectations before promoting it to the next stage.

How to get started

Keptn can be installed through Helm with this simple set of commands:

helm repo add keptn https://charts.lifecycle.keptn.sh
helm repo update
helm upgrade --install keptn keptn/keptn Dynatrace® Software Intelligence Platform
   -n keptn-system --create-namespace –wait

For further information and configuration options, head to the Keptn installation documentation. Otherwise, if your application uses the recommended Kubernetes labels, you are ready to go. Keptn will detect your application and provide observability of your deployment out of the box.

Usually, Keptn does not act in isolation during a deployment. In the logical timeline of events, there will be tools acting before Keptn (such as GitOps tools like ArgoCD) and tools that act after Keptn (such as security scanning tools).

If the tools before Keptn begin generating OpenTelemetry data (for example, spans and traces), it would be beneficial to see Keptn’s portion of the work in the correct context as part of the same distributed trace. Similarly, anything that happens after Keptn but still in the same logical “deployment” operation should be included in the end-to-end trace view.

To achieve this, Keptn now accepts tools to pass W3C trace IDs into Keptn

and Keptn will pass such W3C trace IDs out of Keptn, creating true end-to-end visibility.

With this information, Keptn can track from “PR merged” all the way to “deployment running in production,” allowing DevOps to slice and dice observability data to understand how the application reached that state.

After installing Keptn, you can get started by going through the several guides listed below that will introduce you step by step to the features provided by Keptn.

  • Integrate Keptn with your applications, which shows you how to configure your applications to be deployed with Keptn.
  • OpenTelemetry observability, which shows how to configure Keptn to make use of its observability features, so you can have a holistic overview of your application deployments and any issues that arise during deployment.
  • Multi-stage application delivery, which serves as an example of how you can integrate Keptn with ArgoCD and GitHub to take the previous concepts and apply them across multiple application environments, giving you full observability of your application deployment from development into production.

Conclusion

Keptn empowers DevOps teams to conquer the Kubernetes deployment challenge confidently, ensuring smoother and more efficient deployments. By addressing the complexities inherent in Kubernetes deployments and seamlessly integrating with existing tools, Keptn delivers unparalleled visibility and control. Start your journey with Keptn by downloading the first Release Candidate for v2 available now on GitHub.

The post Mastering Kubernetes deployments with Keptn: a comprehensive guide to enhanced visibility appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/mastering-kubernetes-deployments-with-keptn-a-comprehensive-guide-to-enhanced-visibility/feed/ 0
Enhance data collection with Dynatrace OTel Collector distribution https://www.dynatrace.com/news/blog/enhance-data-collection-with-dynatrace-otel-collector-distribution/ https://www.dynatrace.com/news/blog/enhance-data-collection-with-dynatrace-otel-collector-distribution/#respond Fri, 15 Mar 2024 16:38:36 +0000 https://www.dynatrace.com/news/?p=63069 OpenTelemetry demo app Astronomy Shop

As organizations strive for observability and data democratization, OpenTelemetry emerges as a key technology to create and transfer observability data. OpenTelemetry is gaining popularity because it’s considered a standard, and that’s why it’s a common choice for creating future-proof solutions for years to come. To answer the growing demand for OpenTelemetry, Dynatrace is proud to […]

The post Enhance data collection with Dynatrace OTel Collector distribution appeared first on Dynatrace news.

]]>
OpenTelemetry demo app Astronomy Shop

As organizations strive for observability and data democratization, OpenTelemetry emerges as a key technology to create and transfer observability data. OpenTelemetry is gaining popularity because it’s considered a standard, and that’s why it’s a common choice for creating future-proof solutions for years to come.

To answer the growing demand for OpenTelemetry, Dynatrace is proud to announce the release of the Dynatrace OTel Collector distribution. This collector, fully supported and maintained by Dynatrace, is entirely open source. Before we get into the specifics, let’s first recap the benefits OpenTelemetry offers and why using collectors is a best practice.

Understanding OpenTelemetry

OpenTelemetry is an open, vendor-neutral standard for creating, collecting, and transferring telemetry data, like traces, metrics, and logs. Developers and operators can gain insights into their applications and infrastructure without fear of vendor lock-in because OpenTelemetry is fully open source and owned by CNCF. The OpenTelemetry project is supported and maintained by representatives from Microsoft, Google, Amazon, and many others, including Dynatrace.

Why do I need an OpenTelemetry collector?

As the name suggests, an OpenTelemetry collector gathers data from multiple sources and sends it to observability backends, like Dynatrace, for analysis. A collector helps developers control their telemetry data streams for each signal. Different data streams can be directed to different backends or even multicast to multiple backends simultaneously. The configuration is highly flexible in solving various user needs.

A collector is also a powerful component for data processing. It removes the burden of managing retries, batching, and sampling from monitored applications, which can reduce the CPU and memory requirements of applications. A collector can also transform and enrich the data with additional context. For example, in a Kubernetes environment, a collector can automatically attach metadata about pods and namespaces to all observability data. This ensures that application telemetry is contextualized with the infrastructure, enabling the observability backend to link the application and infrastructure for enhanced insights and root cause analysis.

From a user perspective, a collector can also serve as an open source platform that can be extended with custom components. You can create internal collector components of your own, for example, to receive telemetry data in a special format or to process it in a certain way. Using a collector as a telemetry processing platform can be much easier than creating an entirely new application.

Why the Dynatrace OTel Collector

The OpenTelemetry community releases different distributions of the collector, many of which our customers use to send OpenTelemetry data to Dynatrace. So, why should you consider using the Dynatrace Otel Collector? Quite simply, support and stability.

We have seen many customers identify collectors as a potential solution for their needs. Still, they haven’t been able to deploy collectors in production due to the lack of support. Deploying an open source component without external support and the needed expertise is undoubtedly a risk. That’s why we provide Dynatrace customers with a Dynatrace-supported solution.

The Dynatrace Otel Collector comes with collector components that have been verified by Dynatrace for seamless operation. This removes the burden of manually validating each component and use case. To further help you with your collector journey, we publish configuration examples of typical Dynatrace use cases and best practices to provide a good starting point.

The Dynatrace Otel Collector includes components that we know run stably in production, which means we can offer full Dynatrace support. At the same time, innovations from the OpenTelemetry community can be added to the Dynatrace Otel Collector only after they are mature enough and have proven their stability.

Deployment and the typical use cases

The Dynatrace Otel Collector can be deployed on Kubernetes or Docker using a provided container image or directly on a host with the published binary. For more details, refer to our Dynatrace OTel Collector deployment guide.

In the initial release, the Dynatrace Otel Collector comes with components for:

Additionally, the Dynatrace Otel Collector includes a rich toolset for data processing to enrich, filter, transform, sample, and batch. The complete list of the components is available in our GitHub repository.

Dynatrace OTel Collector diagram with telemetry sources

What’s next

After the initial release, we’ll continue enhancing the Dynatrace OTel Collector with new features and capabilities to make it even easier to integrate with Dynatrace. We also intend to offer more automated methods to deploy the collector with pre-configurations. So, stay tuned.

As a significant contributor to the OpenTelemetry project, Dynatrace remains committed to working with the community and other vendors to enrich its capabilities and make it user-friendly for everyone.

Deploying the Dynatrace Otel Collector takes only minutes and uses the tooling you already know: standalone binary, Docker image, Kubernetes Operator, Helm chart, or a standard manifest file. The configuration maps one-to-one with the collector distributions from the OpenTelemetry community.

The post Enhance data collection with Dynatrace OTel Collector distribution appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/enhance-data-collection-with-dynatrace-otel-collector-distribution/feed/ 0
Progressive delivery done right with feature flags and OpenFeature https://www.dynatrace.com/news/blog/progressive-delivery-done-right/ https://www.dynatrace.com/news/blog/progressive-delivery-done-right/#respond Wed, 31 May 2023 06:57:30 +0000 https://www.dynatrace.com/news/?p=57929 Dynatrace | OpenFeature

“Move fast and break things.” How many times have you heard that? Know what? No one likes that. No one likes breaking things. It’s more like: “Move fast and break things—only because you don’t know how to avoid it.” Granted, as a soundbite, that is nowhere near as catchy. But progressive delivery is essentially the solution […]

The post Progressive delivery done right with feature flags and OpenFeature appeared first on Dynatrace news.

]]>
Dynatrace | OpenFeature

“Move fast and break things.”

How many times have you heard that? Know what? No one likes that. No one likes breaking things. It’s more like: “Move fast and break things—only because you don’t know how to avoid it.

Granted, as a soundbite, that is nowhere near as catchy. But progressive delivery is essentially the solution to that problem: it enables you to move fast and avoid breaking things.

How does an organization, especially if they’re not yet doing continuous integration and continuous delivery (CI/CD), either at all or well, “move fast and not break things”? The answer: Progressive delivery with feature flags and observability.

Progressive delivery encompasses multiple methodologies where DevOps teams introduce new features to small user subsets (or cohorts) slowly or gradually in a controlled manner. By doing so, teams can closely observe new functionality and hopefully roll it out to end users more safely and with fewer errors.

Progressive delivery with feature flagging delivers the following benefits:

  • Decreases risk
  • Provides faster feedback loops
  • Decouples the mechanics of deploying from the act of a release
  • Increases business agility and confidence

Feature flagging is one method of progressive delivery but there are multiple accepted techniques, for example:

  • A/B (or blue/green) releases
  • Canary releases
  • Feature flagging

You can deliver both canary and blue/green deployments without feature flags. Just copy and paste your code, make changes to the “green” version, and deploy “green” alongside “blue.” Then find some way to direct “some” of your users to “green” and leave everyone else on “blue.”

There are problems with this manual approach, though, for example, the following:

  • Manual changes make the tech stack more complex: You also have to manually update routing rules and other changes.
  • It’s also a blunt instrument. It’s hard to be selective with what constitutes “some” users. Usually, it ends up being a percentage-based thing.
  • You need to pay for, manage, and maintain two copies of the application.
  • If you do need to rollback, all new functionality in that release is revoked (even if it still works)

A better way is to organize features and functions using feature flags, which gives you fine-grained control over what to release, when, and to whom.

How feature flagging enables progressive delivery

At their most basic, feature flags are an if/else statement that can alter the application behavior in real-time, at runtime without redeploying an application.

This technical capability brings the following benefits:

  • Realtime application behavior changes without code redeployment
  • No additional infrastructure required: The application is the same; no additional infrastructure or configuration is required
  • Lower cost: No additional infrastructure means no additional cost
  • Faster experimentation: Enabling and disabling functionality with the flip of a switch means you can try out many more versions and hypotheses. You spend time experimenting rather than waiting for CI/CD pipelines to run.

Feature flag use cases

You can use feature flagging in lots of situations, but here are a few examples:

Limiting access: For your eyes only

  • Most start their feature flag journey with scenarios like these:
  • Users in Australia get the new features first, before other countries
  • US users see a localized version of the content
  • Staff get access to additional features
  • Upon login, the “user level” badge is bronze, silver, or gold, depending on the logged-in user
  • A subset of users is selected to receive a special offer. The special offer is placed behind a feature flag. The flag is enabled for those users. After a set amount of time, the flag is disabled – the offer is no longer available.

A “deployment” no longer means a “release”

Using feature flags, “deploying” code no longer means “releasing code to end users.” With feature flags, teams can place new features and functionality behind a feature flag that is disabled by default.

This makes deployments almost risk-free and can dramatically increase your code-to-production time.

After all, if you know that by default, no one gets the new code, your deployment is safe. Then when you’re ready, enabling new functionality for only those you choose means that even if something goes wrong, you have the following safeguards in place:

  • The blast radius is limited; only a small subset of users are impacted
  • You know who it went wrong for and can easily disable the functionality without a redeployment.

Feature flagging with the Dynatrace Platform

The Dynatrace Platform uses feature flags extensively:

  • All new functionality is placed behind feature flags. Customers can opt-in to early access programs. Product managers then enable the flags for those customers only.
  • Dynatrace clusters are feature-flagged as “stable,” “early,” or “developer” to denote a risk appetite. Updates are pushed first to “developer” clusters. If everything is OK, the “early” clusters receive the update until finally, the “stable” clusters are updated.
  • Documentation for new features or “internal only” knowledge is placed behind a feature flag. When logged in, staff have access to documentation that no one else can see.

What if I’m not doing CI/CD yet?

As LaunchDarkly points out, if you’re not doing CI/CD yet, at first glance, feature flagging can seem daunting and “too difficult.”

However, if you can deploy knowing the new code is never seen or used unless intended, why not continuously deliver software?

CI/CD is scary for many precisely because a deployment means a release, and historically, that’s an “all or nothing” activity.

Feature flags give you a safety net, so CI/CD doesn’t have to be scary.

OpenFeature

OpenFeature is an open standard that describes feature flagging. It has already been adopted by many feature flag vendors and adopted by some names you might know.

canonical logo flipt logo
CloudBees logo OpenTelemetry logo

Companies normally start their feature flag journey by building an in-house solution. This could be backed by a database or a simple JSON file.

Sooner or later though, companies invariably find they’ve outgrown (or no longer wish to maintain) their in-house solutions. The business decides on a feature flag vendor, and the developers get to work.

The developers must:

  1. Learn the APIs of the chosen feature flag vendor
  2. Remove all existing integration code between the application(s) and the in-house vendor
  3. Replace all the above code with new code to “speak to” the feature flag vendor
  4. Repeat this process across every application in the enterprise

OpenFeature makes it much simpler for enterprises to adopt feature flags whether they’re using an in-house or commercial flag solution:

  • OpenFeature requests flag values in a standard way, regardless of “backend”
  • Teams can swap backends easily

To enable feature flagging with OpenFeature, all you have to do is change the following value from

OpenFeature.setProvider(“My-In-House-Provider”)

to

OpenFeature.setProvider(“My-Vendor-Provider”)

OpenFeature architecture

The importance of observability

Feature flags and unrestricted continuous delivery bring great benefits, but observability is a critical component of a progressive delivery system. After all, if you can change everything about the system using flags but can’t see what the effects of those changes are, you’re flying blind.

You can configure Dynatrace to capture feature flag values on every distributed trace, so the impacts of those flags can truly be understood at a transaction and individual user level.

Dynatrace makes it easy to see exactly what the issue is and who is impacted.

Example: API traffic with feature flags

Imagine an API endpoint that a service calls to perform an action. This action relies on an algorithm. You’ve enabled feature flags to decide which algorithm a user receives.

Your team wants to introduce and test a new algorithm (which is supposedly faster) and you implement a feature flag to target only a small percentage of logged-in users to test this new algorithm.

To determine whether the new algorithm is actually faster, you must measure its response time and compare it to the response time of the old algorithm. If you only split the response time by endpoint, the statistics for logged out and in would be grouped together – you’d get the average of both algorithms (user groups).

Dynatrace captures the flag values, so splitting by the logged-in or out status is easy.

progressive delivery with feature flags: Capturing flag values with Dynatrace

progressive delivery with feature flags: Capturing flag values with Dynatrace

progressive delivery with feature flags: Capturing flag values with Dynatrace showing who is logged in and logged out

Dynatrace shows that logged-in users have a feature flag enabled which gives a better response time.

Perhaps it is time to roll out the new algorithm to a higher percentage of logged-in users. By using feature flags, you can roll it out slowly and safely. Dynatrace is there every step of the way, observing and ensuring a safe progressive rollout.

progressive delivery with feature flags: Capturing flag values with Dynatrace, determining who is logged in and logged out

Dynatrace ❤️ progressive delivery with feature flags

Start your progressive delivery journey today with Dynatrace. Sign up for a free trial and start realizing safety-guaranteed progressive delivery with feature flags.

When you have your trial tenant, head over to OpenFeature to get started.

To learn more about feature flags with Dynatrace, see my previous post, Feature flagging done right with Dynatrace and OpenFeature.

The post Progressive delivery done right with feature flags and OpenFeature appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/progressive-delivery-done-right/feed/ 0
Feature flags done right with the OpenFeature initiative and Dynatrace https://www.dynatrace.com/news/blog/feature-flags-with-openfeature-and-dynatrace/ https://www.dynatrace.com/news/blog/feature-flags-with-openfeature-and-dynatrace/#respond Tue, 06 Dec 2022 08:57:09 +0000 https://www.dynatrace.com/news/?p=55033 Dynatrace | OpenFeature

Fast becoming a staple practice in CI/CD workflows, feature flags make it easy for developers to turn features on and off like a switch. This gives them the flexibility to test, selectively release, and roll back features on a granular scale without disrupting the mainstream code base. But homegrown efforts result in technical debt and observability issues that require a standardized approach. OpenFeature is the open standard that, coupled with Dynatrace software intelligence, makes it easy to integrate feature flagging into your cloud automation regime.

The post Feature flags done right with the OpenFeature initiative and Dynatrace appeared first on Dynatrace news.

]]>
Dynatrace | OpenFeature

In software development, feature flags are an established path to rapid value, continuous progressive delivery, and safe deployments. The ability to isolate certain software capabilities makes it easier to test, preview, release, and roll back small functional increments. Used by organizations for everything from assigning support tickets to managing failover regimes, feature flags enable DevOps teams to release software faster and more reliably.

But feature flagging can also introduce some issues. In addition to requiring a high degree of custom coding, feature flags can rapidly accrue technical debt that can be opaque to diagnose.

In this post, I’ll explain why Dynatrace and others started the OpenFeature initiative to address these issues and donated it to the Cloud Native Computing Foundation (CNCF) to make feature flagging an open standard.

Prefer to watch than read? Watch “Feature flagging done right with OpenFeature and Dynatrace”



Video thumbnail

What are feature flags?

Feature flags are a software development methodology that enables developers to turn on or off specific functions at runtime. Using scripting tags, feature flags work without having to deploy new code. This approach gives teams a high degree of flexibility and control over testing, beta releases, access privileges, and feature lifecycles.

Feature flags work by building dedicated feature branches into code that make functionality available to certain user groups simultaneously. Also referred to as toggles, feature flags enable development teams to test new features in production with minimal risk. They’re also useful for canary releases, where teams can throw a kill switch to effectively roll back a faulty feature.

Because of their flexibility, feature flags have become a staple practice for CI/CD workflows.

What is the OpenFeature initiative?

The OpenFeature initiative is an open standard CNCF sandbox project for managing feature flags. Submitted by Dynatrace and a consortium of feature flag management solutions, OpenFeature provides a future-proof, vendor-neutral way to integrate feature flagging and management solutions. Because it’s open source, OpenFeature eliminates the need for organizations to build their own proprietary SDKs and APIs. This open source approach to managing feature flags simplifies and accelerates release cycles and enables teams to automate releases at scale.

The problems with feature flags—and their solutions

The most common route into feature flagging starts with a company building its own in-house solution. A homegrown solution meets requirements and works great—until it doesn’t. Inevitably, the company realizes that their solution can’t compete with the established open source or vendor solutions, or they realize they can’t devote the time and resources to maintain their solution.

The company needs and wants to move toward a “better fit” solution, but they run into a couple of problems.

Problem #1: Tight coupling and technical debt

For organizations on the do-it-yourself journey, developers often integrate their applications directly with the in-house feature flag solution. Nothing wrong with that, until they try to move. They have tightly coupled the application logic to the feature flag “vendor” (in this case, their in-house solution).

In other words, their application is now entirely dependent on their in-house solution. They have introduced technical debt, making it costly and time-consuming to migrate to a “better fit” solution.

OpenFeature user to your app to in-house feature flag

Problem #2: Feature flag observability

Companies use feature flags for various reasons, which all boil down to increased business agility, including the following abilities:

  • Enable and disable a feature at the flick of a switch.
  • Deploy code without releasing it to end users such as new functionality hidden by default behind a feature flag.
  • Employ a kill switch, for example, to detect a DDOS attack and automatically enable static content.
  • Experiment and run hypothesis-driven experimentation with real users.

However, knowing that a flag is enabled does not tell the whole story. Teams need crucial context to understand the significance of these data points, such as the following:

  • What impact did that flag have on users?
  • Were conversion goals impacted?
  • Were SRE metrics impacted, such as response time, availability and throughput?

In short, businesses need comprehensive observability of the impact their feature flagging decisions so they can make data and evidence-driven decisions. make data and evidence-driven decisions.

Solution: The OpenFeature initiative

OpenFeature solves these problems. Much like OpenTelemetry, OpenFeature is a vendor-neutral specification rather than a product.

Solution #1: Increase flexibility and remove technical debt

First, the OpenFeature initiative defines a standardized method to interact with feature flag backends. Developers don’t need to worry about the feature flag backend: get_flag_string_value(“name”, “default”) is the same regardless of whether the value is coming from an in-house solution or a vendor solution.

Crucially, if the back-end flag provider needs to change, for example, from in-house to vendor, no technical debt exists, and none of the application code needs to change.

The only component that needs change is the OpenFeature Provider (marked in blue).

OpenFeature user to your app to OpenFeature provider to in-house feature flag DB

OpenFeature user to your app to OpenFeature provider to Vendor backend

Solution #2: Observable by default

OpenFeature comes with OpenTelemetry-based observability by default. OpenTelemetry-compliant observability platforms like Dynatrace can easily consume that data. With comprehensive, AI-driven analysis of observability data, organizations can automatically gain answers to the impact of feature flag statuses and the services they touch. This insight helps teams make critical decisions about what to do and what to automate.

The following Dynatrace dashboard is automatically created by the Dynatrace open source monitoring-as-code standard Monaco as a result of following the hands-on example.

OpenFeature Dashboard

Who writes the feature flags’ translation-layer providers?

Providers are a crucial part of the OpenFeature initiative that translates the OpenFeature “generic” calls and map them to vendor-specific feature flag calls. They’re the middleware, the glue, the translation layer.

The good news is most feature flag vendors have already committed to supporting OpenFeature, so they have already written and maintain providers you can use. No code for you to write. For example, here are the JavaScript providers.

If your feature flag vendor does not yet support the OpenFeature initiative in your language, speak to them about supporting supporting it.

If you need to implement a provider for an in-house solution, the OpenFeature Providers page contains a how-to guide, code samples, and a checklist to get you started.

Explore the OpenFeature initiative further

Here are some ways you can learn more about OpenFeature.

Organizations need to release software at a high velocity to stay competitive as the pace of business accelerates, but they can’t sacrifice software quality for speed. Agile companies that adopt continuous software delivery have increasingly enlisted feature flagging to enable more frequent code releases.

The post Feature flags done right with the OpenFeature initiative and Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/feature-flags-with-openfeature-and-dynatrace/feed/ 0