Cloud native Archives | Dynatrace news https://www.dynatrace.com/news/category/cloud-native/ The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Wed, 08 Jul 2026 06:43:57 +0000 en hourly 1 Getting started with cloud modernization https://www.dynatrace.com/news/blog/getting-started-with-cloud-modernization/ https://www.dynatrace.com/news/blog/getting-started-with-cloud-modernization/#respond Tue, 23 Jun 2026 09:01:32 +0000 https://www.dynatrace.com/news/?p=73774 Gen AI graphic

Here’s the uncomfortable truth: organizations are spending billions on cloud migrations, yet many fail to deliver the promised ROI. Why? Because moving workloads to the cloud without addressing operational visibility, runaway costs, and fragmented tooling multiplies the problem. In this article, we’ll cover: Why most cloud modernization efforts stall (and how to avoid those pitfalls) […]

The post Getting started with cloud modernization appeared first on Dynatrace news.

]]>
Gen AI graphic

Here’s the uncomfortable truth: organizations are spending billions on cloud migrations, yet many fail to deliver the promised ROI. Why? Because moving workloads to the cloud without addressing operational visibility, runaway costs, and fragmented tooling multiplies the problem.

In this article, we’ll cover:

  • Why most cloud modernization efforts stall (and how to avoid those pitfalls)
  • What true cloud modernization looks like when built on the right foundation
  • A phased approach to modernization that delivers measurable business outcomes
  • How to choose the right path forward for your unique applications

Understanding cloud modernization

Cloud modernization refers to the process of updating existing legacy applications and infrastructure to leverage cloud-native technologies, architectures, and methodologies. This transformation goes beyond simple lift and shift migrations to embrace modern approaches like containerization, microservices, and DevOps practices.

The impetus for modernization often comes from challenges with legacy systems:

  • Monolithic architectures that are difficult to update and scale
  • High operational costs and maintenance burdens
  • Inflexibility in meeting changing business requirements
  • Limited ability to integrate with modern services and tools
  • Slow deployment cycles that hamper innovation

But here’s what organizations often don’t anticipate: without unified observability and intelligent automation, cloud modernization often moves these problems to a more complex, distributed environment where they become even harder to diagnose and fix.

When done right, modernization solves these issues and unlocks entirely new capabilities. Teams move from reactive to predictive, applications become self-healing, and innovation accelerates.

Why cloud modernization matters

According to IDC, 83% of enterprises are working to rationalize and optimize their technology infrastructure, yet only 35% consider their approach to modernization effective. This gap highlights the need for thoughtful, strategic approaches to cloud modernization.

The benefits of a well-executed cloud modernization strategy include:

  • Enhanced application insights: Observability solutions provide deeper visibility into application performance, dependencies, and user activity, enabling proactive optimization.
  • Improved performance: Cloud-native architectures and services can dramatically improve application responsiveness and reliability.
  • Cost optimization: Modernized applications typically reduce operational costs through more efficient resource utilization and automated management.
  • Increased operational efficiency: DevOps practices and automation reduce manual intervention, allowing IT teams to focus on innovation rather than maintenance.
  • Business agility: Perhaps most importantly, modernized applications can adapt quickly to changing market conditions and business requirements.

Essential elements of a cloud modernization strategy

Assessment and planning

Before embarking on a modernization journey, organizations should conduct a thorough assessment of their existing application portfolio. This includes cataloging all applications and their interdependencies, evaluating each application’s business value and technical debt, understanding current performance metrics and cost structures, and identifying modernization priorities based on business impact.

For example, a financial services company might prioritize modernizing its customer-facing applications to improve user experience while planning a phased approach for back-office systems.

Choosing the right modernization approach

Not all applications require the same modernization approach. The key is matching your strategy to business impact and technical reality.

Ask three key questions to guide your decision:

  1. What’s the business value? Customer-facing applications typically warrant more investment than back-office systems.
  2. What’s the technical debt? Minor issues suggest lighter modernization while fundamental problems require deeper transformation.
  3. What’s the integration complexity? Standalone applications are easier to modernize than heavily interdependent systems.

Based on your answers, you’ll typically follow one of three paths:

When to use What it means Example
Path 1: Optimize in place (rehosting or replatforming) Low technical debt, need cloud benefits quickly Move to cloud with minimal or moderate changes to leverage capabilities like auto-scaling and managed services A financial services company replatforms its customer-facing mobile app to take advantage of cloud auto-scaling and global distribution
Path 2: Transform the architecture (refactoring or rearchitecting) High-value applications with significant technical debt Restructure existing code or fundamentally redesign how the application works to become cloud-native A manufacturing company rearchitects its supply chain analytics platform to use microservices, enabling independent scaling and faster deployment
Path 3: Replace entirely (rebuilding or switching to SaaS) Commercial solutions exist or complete redesign delivers clear ROI Completely redesign and rewrite the application for the cloud, or switch to proven SaaS alternatives A retail organization replaces its custom CRM while rebuilding its inventory optimization engine to leverage cloud-native machine learning services

The specific tactics within each path – whether you’re rehosting, replatforming, refactoring, rearchitecting, rebuilding, or replacing – depend on your application’s complexity, value, and business requirements.

Accelerating software delivery

Cloud modernization should enhance an organization’s ability to deliver software quickly and reliably. This involves implementing CI/CD pipelines for automated testing and deployment, using infrastructure as code for consistent environment provisioning, leveraging automated quality and security gates, and adopting platform engineering approaches to provide self-service capabilities for developers.

Increasing team productivity

The shift-left/shift-right approach to cloud modernization focuses on sharing responsibility between development and operations teams while maintaining centralized governance. This hybrid model provides developers with self-service capabilities, centralizes expert knowledge within platform engineering teams, maintains consistent tooling and knowledge management, and balances autonomy with governance.

Getting started with cloud modernization

To begin your cloud modernization journey, successful organizations often break the process into three clear phases:

Phase 1: Establish your foundation (months 1 to 3)

  1. Define measurable outcomes such as reduced time-to-market, improved user satisfaction, or decreased operational costs.
  2. Benchmark existing applications, processes, and services to create a baseline for measuring improvement.
  3. Choose a pilot application that provides real business value but isn’t mission-critical, allowing your team to learn and refine approaches before tackling more complex systems.
  4. Implement comprehensive monitoring and observability solutions early in your modernization process to gain insights that inform future decisions.
  5. Success milestone: Complete visibility into your environment, pilot selected, and baseline established

Phase 2: Modernize and automate (months 4 to 9)

  • Identify repetitive tasks and implement automation, focusing on areas that will yield the highest return on investment. Build CI/CD pipelines and automated quality gates.
  • Enable self-service for development teams while maintaining centralized governance through platform engineering.
  • Success milestone: Pilot in production with measurable improvements; automation reducing manual effort by 50% or more

Phase 3: Scale and optimize (months 10 and beyond)

  • Expand to additional applications using lessons learned from your pilot.
  • Leverage AI-powered insights for continuous optimization of costs, performance, and reliability.
  • Achieve operational maturity where routine operations are automated and teams focus on innovation.
  • Success milestone: Measurable ROI across applications; 40% increase in team productivity; 30% reduction in operational costs

Accelerate cloud modernization with Dynatrace

Modernizing your cloud environment helps to unlock agility, resilience, and long-term business value.

With Dynatrace, you get unified observability, precise automation, and advanced AI capabilities that help you modernize with confidence. From predictive issue detection to automatic root-cause analysis and intelligent remediation, Dynatrace enables you to scale operations, accelerate innovation, and reduce complexity—all while maintaining control and compliance.

What makes Dynatrace different:

  • Unified data model: One source of truth across your entire stack eliminates tool sprawl and conflicting data. Every team works from the same real-time information, speeding decisions and collaboration.
  • Causal automation: Root cause analysis happens in real-time. Dynatrace Intelligence doesn’t just detect anomalies—it automatically determines why they happened and triggers intelligent remediation capabilities through agentic AI. Teams shift from firefighting to innovating.
  • Continuous topology awareness: As your environment changes constantly, Dynatrace automatically discovers, maps, and understands every component and dependency in real time. No manual configuration. No blind spots. You always know what’s running and how it’s connected.

Take the next step

Zurich North America cut IT service incidents by 89% while accelerating its cloud migration with Dynatrace. See how, then start your own journey with a free trial.

Or take Dynatrace for a spin by exploring the public sandbox.

Cloud modernization: Frequently asked questions

What is cloud modernization?

Cloud modernization is the process of evolving legacy applications and infrastructure to take advantage of cloud-native architectures, services, and operating models. It goes beyond migration to improve agility, scalability, and efficiency.

How is cloud modernization different from cloud migration?

Cloud migration focuses on moving workloads to the cloud. Cloud modernization focuses on improving how applications are built, run, and operated once they’re there—often through automation, observability, and architectural change.

What are the main approaches to cloud modernization?

Organizations typically modernize by optimizing in place (rehosting or replatforming), transforming the architecture (refactoring or rearchitecting), or replacing applications entirely with rebuilt or SaaS solutions.

How does AI accelerate cloud modernization?

AI turns observability data into actionable answers, helping teams understand dependencies, prioritize actions, and automate operations with confidence. Dynatrace AI capabilities, powered by Dynatrace Intelligence, automatically detect anomalies, identify root causes, and prioritize remediation, so teams can modernize faster without relying on manual investigation.

How do I choose the right modernization approach for my applications?

The right approach depends on business value, technical debt, and integration complexity. High-value, customer-facing applications often justify deeper modernization, while lower-risk systems may benefit from lighter optimization.

Why is observability important in cloud modernization?

Observability provides real-time insight into application performance, dependencies, and costs. It helps teams make informed modernization decisions, detect issues early, and continuously optimize as environments change.

What’s the best way to get started with cloud modernization?

Start by assessing your application landscape, defining measurable business goals, and selecting a pilot application. Establish visibility early, then modernize in phases to prove value before scaling.

The post Getting started with cloud modernization appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/getting-started-with-cloud-modernization/feed/ 0
Want to catch the AI native wave? Learn the lessons of the cloud-native shift https://www.dynatrace.com/news/blog/want-to-catch-the-ai-native-wave-learn-the-lessons-of-the-cloud-native-shift/ https://www.dynatrace.com/news/blog/want-to-catch-the-ai-native-wave-learn-the-lessons-of-the-cloud-native-shift/#respond Mon, 24 Nov 2025 18:35:51 +0000 https://www.dynatrace.com/news/?p=72006 PurePerformance podcast

From the personal computer to the internet to mobile and cloud native, transformative technologies bring new challenges and opportunities. While every innovation is different, Pini Reznik says there are patterns that can be applied when adopting any new technology. In the latest PurePerformance podcast, Reznik talks to hosts Andi Grabner and Brian Wilson about his […]

The post Want to catch the AI native wave? Learn the lessons of the cloud-native shift appeared first on Dynatrace news.

]]>
PurePerformance podcast

From the personal computer to the internet to mobile and cloud native, transformative technologies bring new challenges and opportunities. While every innovation is different, Pini Reznik says there are patterns that can be applied when adopting any new technology.

In the latest PurePerformance podcast, Reznik talks to hosts Andi Grabner and Brian Wilson about his new book From Cloud Native to AI Native: Catching the Next Wave of Innovation, and how organizations can take a pragmatic approach to AI adoption.

The AI Native transformation process

Here’s how Reznik describes the process of adoption AI, or any other transformational technology:

  • Experiment: Start with a small, skunkworks-style team operating in a sandbox. Their mission isn’t to deliver production-ready systems but to experiment, learn, and identify viable business cases.
  • Find a small win: Once you find a promising use case, build a minimal viable product that demonstrates tangible value. Success here justifies incremental investment.
  • Build a foundation: Build the infrastructure, develop the team structure, and foster the new culture necessary to adopt the technology at scale.
  • Scale up: Expand as you grow, transitioning from legacy systems to new solutions carefully.

The biggest mistake organizations make is skipping the first two phases and jumping straight to large-scale initiatives. That’s why you hear so much about failed AI pilots: Organizations try to scale before they validate their use cases.

“Transformation isn’t a one-time project; It’s a way of thinking.”

— Pini Reznik, CEO and Co-Founder, re:cinq

AI for the little guy

In the previous episode of PurePerformance, Laura Tacho made the case for AI’s true potential in software development being not in code generation but in speeding up feedback loops and helping ensure that developers build the right things. Where Tacho focused on developer experience and the software development lifecycle, Reznik focuses on organization-wide potential, particularly for companies that haven’t traditionally built much, if any, software in-house. AI could make it affordable for smaller companies to build custom software based around their unique value propositions.

Our perspective: Start small, validate value early, and ensure every step is measurable and observable. When teams can see the impact of AI in real time—on performance, cost, and outcomes—they make smarter decisions and scale with purpose.

To to dive into the world of software performance and innovation, listen to the latest episode of PurePerformance.

The post Want to catch the AI native wave? Learn the lessons of the cloud-native shift appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/want-to-catch-the-ai-native-wave-learn-the-lessons-of-the-cloud-native-shift/feed/ 0
Transforming Azure Data Factory operations with Dynatrace https://www.dynatrace.com/news/blog/transforming-azure-data-factory-operations-with-dynatrace/ https://www.dynatrace.com/news/blog/transforming-azure-data-factory-operations-with-dynatrace/#respond Tue, 09 Sep 2025 15:01:53 +0000 https://www.dynatrace.com/news/?p=70939 Alerting and Analyzing

Data pipelines are critical to seamless operations and informed decision making in modern businesses, and efficiently managing and monitoring those pipelines is crucial for maintaining a competitive edge. Without visibility into pipeline performance, teams risk delays, data loss, and costly downtime. As data volumes grow and workflows become more complex, the stakes get higher—making intelligent […]

The post Transforming Azure Data Factory operations with Dynatrace appeared first on Dynatrace news.

]]>
Alerting and Analyzing

Data pipelines are critical to seamless operations and informed decision making in modern businesses, and efficiently managing and monitoring those pipelines is crucial for maintaining a competitive edge. Without visibility into pipeline performance, teams risk delays, data loss, and costly downtime. As data volumes grow and workflows become more complex, the stakes get higher—making intelligent observability a must-have, not a nice-to-have.

Azure Data Factory (ADF) is a powerful tool for orchestrating and automating data workflows. Whether you’re moving data across hybrid environments, transforming it for analytics, or syncing it between systems, ADF provides the flexibility and scalability needed for modern data operations. Data engineers who deploy Dynatrace with ADF gain valuable insights into performance, business analytics, and automation.

Identifying and resolving pipeline bottlenecks

Keeping data pipelines fast, reliable, and scalable is a huge challenge. When performance dips or failures occur, it’s often a scramble to pinpoint the issue. With Dynatrace, you gain clear visibility into pipeline behavior, resource usage, and failure patterns—making troubleshooting faster and optimization smarter.

Dynatrace addresses key questions such as:

  • What are my longest running pipelines?
    Identifying pipelines with extended runtimes helps teams spot inefficiencies in data processing or transformation logic so that you can optimize performance and reduce latency in downstream systems.
  • Do I have any failing pipelines?
    Immediate visibility into failures allows teams to respond quickly, minimizing data loss and avoiding disruptions to business-critical workflows.
  • Why are my pipelines failing?
    Understanding the root cause—whether it’s a misconfigured activity, resource constraint, or external dependency—enables faster resolution and helps prevent repeat incidents.
  • Which pipelines require optimization?
    By highlighting pipelines with high resource consumption or inconsistent performance, Dynatrace helps prioritize tuning efforts for maximum impact.
  • Do I need to scale resources or adjust concurrency settings?
    These insights guide infrastructure decisions, ensuring that pipelines run efficiently without overprovisioning or underutilizing resources.

By ingesting logs and metrics from Azure Monitor and correlating diagnostics from Azure Data Factory, Dynatrace delivers actionable insights and a comprehensive view of pipeline performance.

Azure Data Factory dashboard in Dynatrace screenshot

Azure Data Factory metrics dashboard in Dynatrace screenshot

Let’s take a closer look at the dashboard. In addition to status and duration, the captured logs and metrics allow us to thoroughly analyze ADF performance.

Azure Data Factory captured logs and metrics in Dynatrace screenshot

  • Time spent in queue: Indicates how long the pipeline waits before starting. Prolonged queue times might suggest a need to scale resources, improve scheduling, or adjust concurrency settings.
  • Time spent in progress: Reflects the actual execution time of the pipeline. Longer durations here may highlight opportunities to optimize pipeline logic or resource allocation.
  • Message: Displays detailed runtime information, including errors. Expanding this field provides additional insights for troubleshooting.

Azure Data Factory message in Dynatrace screenshot

Effective pipeline monitoring transforms reactive troubleshooting into proactive optimization. Leveraging these insights, organizations can not only resolve current bottlenecks but also establish best practices for sustainable, high-performing data environments.

Achieving real-time business insights

Unlocking real-time business analytics isn’t just about tracking technical metrics—it’s about connecting IT operations to business outcomes. With Azure Data Factory (ADF) and Dynatrace, practitioners can enrich pipeline observability by embedding business context directly into logs and metrics. This allows teams to monitor not just how pipelines are running, but what they’re delivering.

For example, imagine a pipeline processing a spreadsheet containing daily revenue figures. By using a Lookup activity in ADF, you can extract key business values—like total revenue or transaction count—and pass them as custom user properties into Dynatrace. This enables dashboards that show not only pipeline health, but also business impact: Was revenue successfully ingested today? Did a failure affect a critical report?

This integration empowers teams to:

  • Track business KPIs alongside technical metrics, making it easier to prioritize fixes based on impact.
  • Spot anomalies in business data early, such as missing values or unexpected drops in volume.
  • Align IT and business teams, fostering collaboration through shared visibility into what matters most.

By bridging the gap between data operations and business insights, Dynatrace helps practitioners move from reactive monitoring to strategic decision-making.

Dynatrace integration in Azure Data Factory

ADF Business Analytics metric in dashboard in Dynatrace screenshot

Ensuring pipeline reliability with automation

Harness the power of automation with Dynatrace Workflows, allowing you to streamline processes with Azure Data Factory. With the ADF REST API, you can configure automation to meet your unique business needs.

Consider the case where an administrator needed a way to automatically retry pipelines on failure for those managed by different teams. While native retry policies were available, this simple automation ensured that retries were executed even when the configuration was overlooked.

This is a perfect example of shifting from reactive troubleshooting to proactive reliability. Instead of waiting for failures to be manually addressed, Dynatrace enables automated responses that reduce downtime, improve consistency, and free up teams to focus on higher-value work. By embedding automation into pipeline operations, practitioners can build more resilient systems and ensure that critical workflows stay on track—even when things go wrong.

Azure Data Factory workflow in Dynatrace screenshot

Get started

By integrating Dynatrace with Azure Data Factory, practitioners gain more than just monitoring: they unlock a smarter, more proactive way to manage data pipelines. From identifying bottlenecks and failures to embedding business context and automating recovery, Dynatrace transforms pipeline operations into a strategic advantage.

The key benefits:

  • Faster troubleshooting with deep visibility into pipeline behavior and failure patterns
  • Smarter optimization through performance metrics and resource insights
  • Real-time business analytics by linking IT operations to business outcomes
  • Proactive reliability with automated workflows that reduce downtime and manual effort

Modern cloud environments need an expanded approach to observability. Learn more about how Dynatrace can help you say goodbye to cloud complexity. Or explore it for yourself in our public sandbox environment.

The post Transforming Azure Data Factory operations with Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/transforming-azure-data-factory-operations-with-dynatrace/feed/ 0
Maximizing cloud efficiency: Driving cost optimization and sustainability with Dynatrace https://www.dynatrace.com/news/blog/optimize-cloud-cost-cloud-sustainability-with-dynatrace/ https://www.dynatrace.com/news/blog/optimize-cloud-cost-cloud-sustainability-with-dynatrace/#respond Thu, 26 Jun 2025 19:39:21 +0000 https://www.dynatrace.com/news/?p=69592 Dynatrace for Executives: Cloud cost optimization & sustainability

As organizations embrace a cloud- and AI-native future, the pressure to control infrastructure spending while meeting sustainability goals intensifies. As a CTO, I want my investments to go into people—building strong, innovative development teams—rather than overspending on cloud resources that don’t deliver business value. This is where Dynatrace plays a crucial role: helping organizations optimize […]

The post Maximizing cloud efficiency: Driving cost optimization and sustainability with Dynatrace appeared first on Dynatrace news.

]]>
Dynatrace for Executives: Cloud cost optimization & sustainability

As organizations embrace a cloud- and AI-native future, the pressure to control infrastructure spending while meeting sustainability goals intensifies. As a CTO, I want my investments to go into people—building strong, innovative development teams—rather than overspending on cloud resources that don’t deliver business value.

This is where Dynatrace plays a crucial role: helping organizations optimize cloud costs while advancing sustainability goals and enabling AI innovation. These priorities are no longer at odds; instead, they go hand in hand.

Key insights for executives

  • AI innovation and sustainability goals can go hand in hand. As AI workloads surge—projected to exceed 50% of cloud compute by 2028—organizations must balance innovation with cost and environmental impact. Dynatrace enables both by optimizing cloud usage in real time.
  • AI-powered observability empowers teams to cut waste, reduce emissions and align spending with business value. By providing deep insights into idle resources, inefficient architecture, and energy-heavy workloads.
  • Dynatrace improves efficiency and supports sustainability goals by dynamically scaling resources based on real-time data demand and business goals.

The true cost of cloud and AI

The rapid spread of AI—including LLMs, agentic AI systems, and coding assistants—and the shift to dynamic, multicloud environments have created yet more layers of complexity. AI workloads are compute-intensive. In fact, one of our customers in the banking sector shared that GenAI tasks cost five times more than traditional cloud workloads. Gartner® predicts that by 2028, more than 50% of cloud compute will be AI-related, up from just 10% in 2023.1

The International Energy Agency – Electricity 2024 report stated that when comparing the average electricity demand of a typical Google search (0.3 Wh of electricity) to OpenAI’s ChatGPT (2.9 Wh per request), and considering 9 billion searches daily, this would require almost 10 TWh of additional electricity in a year. That’s enough to power approximately 3 million households—or all private households in London—with energy for a year.

This growth brings significant environmental and financial implications. Yet, most organizations are beholden to opaque carbon footprint multipliers calculated by the cloud provider, which is often insufficient for actionable insights. Similarly, traditional cost reporting tools lack the depth and runtime visibility needed to drive meaningful optimization.

Dynatrace fills that gap, combining real-time observability, AI-powered insights, and topology-aware mapping to bring deep clarity into both cost and carbon impact.

Four steps to smarter cloud cost and energy management

1. Eliminate waste from idle or underutilized resources

Much of today’s cloud waste stems from overprovisioning and forgotten instances, especially in development and AI workloads. Dynatrace automatically detects underutilized or idle resources across the environments and surfaces insights that can drive decisions whether to shut down or re-size them, reducing both spend and carbon footprint.

Smartscape® automatic discovery and topology mapping adds unique value here, showing not just what’s idle, but whether it’s tied to business-critical processes or genuinely redundant.

2. Align cloud consumption to business value

Executives need more than cost data—they need to understand the why behind consumption. Dynatrace connects cloud utilization directly to applications, users, and business processes, enabling teams to assess whether resources are delivering real business value.

By linking costs to outcomes, organizations can prioritize what to keep, right-size what’s inefficient, and decommission what’s no longer serving a purpose.

3. Optimize architecture and energy efficiency

Most organizations are already taking basic steps like using contracted discount reserved instances or more flexible on-demand spot instances. The next level is architectural and source code optimization, such as green architecture and green coding. Dynatrace helps identify inefficient data flows, underperforming services, and high-cost cross-region transfers.

These insights enable teams to apply green coding techniques, reduce energy-hungry compute patterns, and bring data flows closer to where they’re needed, cutting both cost and carbon emissions.

4. Enable smart, automated orchestration

Finally, Dynatrace has a clear vision to make operations more autonomous. Its predictive, AI-driven orchestration of cloud resources enables teams to automatically scale resources up or down based on real-time demand, user behavior, and business impact.

However, autoscaling based on cloud metrics alone can’t ensure a great user experience or cost efficiency. Dynatrace links infrastructure and deep application observability to user-facing outcomes, allowing for smarter scaling that adapts dynamically to seasonal spikes, new product launches, or unexpected load while eliminating idle time and energy waste.

Accelerating sustainable innovation

Sustainability is now a strategic lever, not just a compliance checkbox. It resonates with environmentally conscious customers and a new generation of employees who want to work for conscientious companies.

By using Dynatrace Cost & Carbon Optimization and full-stack observability, organizations can:

  • Gain real-time, fine-grained insights into the energy and carbon impact of workloads
  • Make carbon reporting actionable and automatable instead of superficial
  • Build a more efficient, resilient, and future-proof cloud environment
Dynatrace Carbon Impact & Optimization dashboard
Figure1: Dynatrace Cost & Carbon Impact homepage

Imagine your cloud-native teams rapidly scaling up environments to test the scalability of new AI features, leading to a 40% spike in compute usage. Without visibility, one wouldn’t notice that this test left over idle or oversized instances, quietly driving up both cloud costs and carbon emissions. Now imagine having real-time insights from Dynatrace that reveal 200 idle instances across non-critical environments, costing thousands monthly and consuming unnecessary energy. Dynatrace AI leverages Smartscape® real-time topology to know automatically which instances can be confidently decommissioned or right-sized—cutting waste, aligning spend to business value, and advancing your sustainability goals.

The bottom line: Intelligent clouds mean a more sustainable planet

Organizations today must move beyond basic FinOps or simple sustainability checklists. The future lies in intelligent, self-optimizing clouds that balance performance, cost, and sustainability in real time.

Dynatrace empowers executives to realize this vision—transforming cloud environments into engines of innovation that are efficient, responsible, and aligned with business and environmental goals.

Follow along the new “Dynatrace for Executives” blog series. I’m diving deeper into each of the nine executive use case areas to help you unlock the potential of Dynatrace.
Want to learn more about all nine use cases? See the overview on the homepage.

1 Gartner Press Release, “Gartner IT Symposium/Xpo 2024 Orlando: Day 3 Highlights,” October 23, 2024, https://www.gartner.com/en/newsroom/press-releases/2024-10-23-gartner-it-symposium-xpo-2024-orlando-day-3-highlights.

GARTNER is a registered trademark and service mark of Gartner, Inc. and/or its affiliates in the U.S. and internationally and is used herein with permission. All rights reserved.

The post Maximizing cloud efficiency: Driving cost optimization and sustainability with Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/optimize-cloud-cost-cloud-sustainability-with-dynatrace/feed/ 0
Enhance efficiency and compliance with automated AWS tag change triggers: A step-by-step guide https://www.dynatrace.com/news/blog/srg-aws-tag-changes/ https://www.dynatrace.com/news/blog/srg-aws-tag-changes/#respond Wed, 02 Apr 2025 15:52:01 +0000 https://www.dynatrace.com/news/?p=68498 Site Reliability Guardian

Streamlining site reliability at scale can be daunting, particularly with large-scale AWS environments and architecture that rely on hundreds—or even thousands—of Amazon EC2 instances. However, you can simplify the process by automating guardians in the Site Reliability Guardian (SRG) to trigger whenever there are AWS tag changes, helping teams improve compliance and effectively manage system […]

The post Enhance efficiency and compliance with automated AWS tag change triggers: A step-by-step guide appeared first on Dynatrace news.

]]>
Site Reliability Guardian

Streamlining site reliability at scale can be daunting, particularly with large-scale AWS environments and architecture that rely on hundreds—or even thousands—of Amazon EC2 instances. However, you can simplify the process by automating guardians in the Site Reliability Guardian (SRG) to trigger whenever there are AWS tag changes, helping teams improve compliance and effectively manage system performance.

This step-by-step guide will show you how to configure your architecture to trigger guardians whenever EC2 tags are updated. Note that EC2 is an example; this guide can be made to work generically for tag changes on any AWS resource. By the end of this guide, you’ll be ready to automate guardians at scale and optimize Amazon EC2 management with ease.

Why automate guardians for AWS tag changes?

Before diving into the technical setup, here’s why automating guardians whenever EC2 tags change is beneficial for your organization:

  • Greater efficiency: Automatically triggering guardians removes the need for manual intervention, saving time for DevOps or site reliability engineering (SRE) teams and allowing for more efficient resource management at scale.
  • Better compliance: Automating guardians ensures critical policies and checks are consistently applied after changes across your architecture, improving security and compliance efforts.
  • Cost optimization: Immediate responses to tag changes lead to informed decisions about scaling, shutting down unused instances, or fine-tuning resource efficiency.
  • Proactive site reliability: Automated guardians can monitor the four golden signals, enabling proactive reliability measures.

Now, let’s get started with the setup!

Step 1: Create an API token

Step 1: Create an API token

First, create an API token to integrate AWS services with Dynatrace for guardian automation.

  1. Log into your Dynatrace tenant

Log in to your Dynatrace tenant and note the first part of the URL (for instance, “abc12345”), which is your tenant ID.

  1. Access token settings

Press Ctrl + K or CMD + K and search for “Access Tokens” within Dynatrace.

  1. Generate a new token

Create a new access token and assign it “bizevents.ingest” permissions.

  1. Save the token

Copy and securely store the token, which looks like “dt0c01.*****.*****”. You’ll use this later during configuration.

Step 2: Create the EventBridge connection

Create the EventBridge connection

Configure invocation

Create the EventBridge connection

Amazon EventBridge acts as the bridge between AWS and Dynatrace. Here’s how to set it up:

  1. Navigate to Amazon EventBridge

Log in to your AWS Management Console and go to EventBridge > Connections.

  1. Recreate the cURL command

You can use this cURL command as a reference to establish your connection:

curl -X POST \
'https://abc12345.live.dynatrace.com/api/v2/bizevents/ingest' \
-H 'Authorization: Api-Token dt0c01.*****.*****' \
-H 'Content-Type: application/cloudevent+json' \
-d '{…}'
  1. Set the Authorization Method

Create a new EventBridge connection with the Authorization Method set to “API Key” and use the API token from Step 1 as the value (i.e., “Api-Token dt0c01.*****.****”).

Reminder: The API token is a sensitive value and should be stored in an encrypted format using a tool like AWS Secrets Manager. When following this guide, AWS Secrets Manager is already used.

Step 3: Define and configure EventBridge rules

Event pattern

EventBridge rules define the exact conditions for triggering guardians:

  1. Specify the input template:

Create an input template to modify your event data:

{
  "specversion": "1.0",
  "id": "<id>",
  "source": "aws.<source>",
  "type": "ec2.tag.change",
  "time": "<time>",
  "aws.region": "<region>",
  "aws.eventbridge.rule.arn": "<aws.events.rule-arn>",
  "aws.resources": <resources>,
  "data": <detail>
}
  1. Set the event pattern

Create a rule in EventBridge with the following event pattern:

{
  "source": ["aws.tag"],
  "detail-type": ["Tag Change on Resource"],
  "detail": {
    "service": ["ec2"],
    "resource-type": ["instance"]
  }
}
  1. Apply targets and permissions

Apply targets and permissions

Apply targets and permissions

Assign targets and permissions to ensure successful data ingestion into Dynatrace. Use an IAM role to permit EventBridge to call your API destination.

Input transformer

The input transformer should be set as follows:

{"detail":"$.detail","id":"$.id","region":"$.region","resources":"$.resources","source":"$.source","time":"$.time"}

Step 4: Test tag changes on Amazon EC2 instances

To validate your configuration, perform the following:

  1. Change a tag

Modify a tag by going to your Amazon EC2 instances in the AWS Management Console. For instance, update the “Environment” tag with a new value.

  1. Verify event logging

Check the EventBridge console to ensure your tag change triggered the appropriate event.

  1. Confirm data in Dynatrace

Within Dynatrace, press CMD/Ctrl + K and search for “Notebooks.” Create a new notebook and run the following query:

fetch bizevents | filter event.type == "ec2.tag.change"

If the query returns results, your configuration is working correctly.

Test tag changes on Amazon EC2 instances

Step 5: Set up the guardian

  1. Create a new guardian

Set up the guardian

In Dynatrace, search for “Site Reliability Guardian” (`CMD/Ctrl + K`) and create a new guardian. For best practices, use the “Four Golden Signals” template.

  1. Automate the workflow

Set up the guardian

Either on the overview page showing all guardians or on the analysis page of a selected guardian, click the Automate button. This will generate a workflow that triggers the guardian based on incoming bizevents (Business events). Configure the event type as `bizevent` and set the filter query to:

event.type == "ec2.tag.change"

Set up the guardian

  1. Add a pause

Set up the guardian

To allow your systems to stabilize before triggering the guardian, add a “wait before” step. For example, set a delay of 600 seconds (10 minutes).

Example timeline:

  • 06:59: Tag changed on EC2 instance.
  • 07:00: EventBridge triggers the workflow.
  • 07:10: Guardian is executed after 10-minute pause.
  1. Save the workflow

Save your final workflow to activate the automation.

Step 6: Validate and monitor the setup

Perform end-to-end validation by changing an EC2 tag again. Confirm the following:

  • The tag change event reaches Dynatrace.
  • The workflow triggers the guardian.
  • The guardian results appear in Dynatrace (e.g., heatmaps or relevant logs).

Run the following query in Dynatrace for additional monitoring:

fetch bizevents | filter event.type == "ec2.tag.change"

You should see log entries confirming the successful execution of your guardian process.

Achieve more with Site Reliability Guardian

In this blog, we highlighted the significant benefits of automating Site Reliability Guardian  triggers for Amazon EC2 changes. With automation, SRG helps engineering teams achieve efficiency, improved compliance, and cost optimization.

Learn more about Site Reliability Guardian in our documentation page.
Looking for more insights and support? Join the Automation Guild.

The post Enhance efficiency and compliance with automated AWS tag change triggers: A step-by-step guide appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/srg-aws-tag-changes/feed/ 0
Observability as Code: DIY with Crossplane https://www.dynatrace.com/news/blog/observability-as-code-diy-with-crossplane/ https://www.dynatrace.com/news/blog/observability-as-code-diy-with-crossplane/#respond Fri, 14 Feb 2025 15:16:49 +0000 https://www.dynatrace.com/news/?p=67882 Observability as code graphic

In today’s rapidly evolving cloud-native landscape, managing observability efficiently can be a game-changer for any platform. During our recent tech talk at #KCDAustria, Observability as Code – DIY with Crossplane, we demonstrated how the Upjet project can streamline the creation of a custom Crossplane Dynatrace provider, unlocking robust observability capabilities in Kubernetes environments. This blog […]

The post Observability as Code: DIY with Crossplane appeared first on Dynatrace news.

]]>
Observability as code graphic

In today’s rapidly evolving cloud-native landscape, managing observability efficiently can be a game-changer for any platform. During our recent tech talk at #KCDAustria, Observability as Code – DIY with Crossplane, we demonstrated how the Upjet project can streamline the creation of a custom Crossplane Dynatrace provider, unlocking robust observability capabilities in Kubernetes environments. This blog post covers what we did and how we did it.

Crossplane and Dynatrace

In this blog post, we will deploy a monitoring dashboard with alerts and notifications as one simple Kubernetes resource. To help us achieve our goal, we have to set up a Kubernetes cluster and infrastructure components. But before that, we need to discuss some important concepts and patterns.

Configuration as Code

Managing vast amounts of configurations for organizational setups at scale is a hard problem to solve since they span over many tools and providers and have many different contributors.

However, there is a pattern for remedying many of these problems: treating these configurations as declarative code instead of applying changes manually.

Originally dubbed Infrastructure as Code, the pattern can be generalized and used for anything that provides a proper interface—simply put, Configuration as Code.

While many tools and their respective approaches exist to write and apply such configuration, one of the most interesting recent developments is extending Kubernetes and using its readily available REST API and reconciliation loops.

The operator pattern in Kubernetes

The mechanism that Kubernetes provides for interface extension is called the operator pattern. This is a powerful mechanism for automating the management of complex applications. It extends the Kubernetes API via custom resource definitions (CRDs), enabling new object types to be created. These are constantly watched by custom controllers, so-called operators.

Operators are designed with a “reconciliation loop,” meaning they continuously compare a resource’s real state against its desired state. When a resource deviates, the operator brings it back into alignment. This is the essence of Kubernetes automation and declarative infrastructure.

Crossplane and the operator pattern

Crossplane builds on the operator pattern and extends Kubernetes beyond managing containerized workloads. Through the Kubernetes API, it enables you to define and provision—among other things—cloud infrastructure resources such as databases, compute instances, and networking components.

Crossplane providers implement the operator pattern for external systems (for example, AWS, GCP, Azure). When you install one, it installs its external resources as Kubernetes native custom resource definitions (CRDs). The provider’s controller watches for changes in the desired state of objects—instanced from these CRDs, reconciling them to ensure they match your expectations.

Compositions

In Kubernetes, low-level resources are managed by high-level resources. Crossplane also allows you to build high-level resources using the Composition pattern.

An example would be an Application resource that abstracts away details like database, network, and compute needs. A Crossplane composition enables you to build something like this, giving you control over the new interface (also a Kubernetes CRD) and the implementation (which low-level resources are created and how they are created).

Using the Upjet project

Now that we have an overview of the concepts, let’s look at how we implement our demo. When building a custom Crossplane provider, we can take two approaches: build the provider from scratch or leverage an existing tool. We opted for the latter, by using the Upjet project, which automates the creation of Crossplane providers based on existing Terraform providers. Here’s why:

  1. Speed and simplicity: Writing a provider from scratch requires a deep understanding of the external system API and how Crossplane manages resources. Upjet allows us to generate a provider much faster by transforming Terraform provider schemas into Crossplane CRDs, significantly reducing the development time.
  2. Reuse of Terraform providers: Upjet allows us to tap into the vast ecosystem of Terraform providers. Since there are already many well-established Terraform providers for various cloud platforms and services, using Upjet means we don’t have to reinvent the wheel.

Upjet project diagram with Crossplane and Dynatrace

Generating a new Crossplane provider with Upjet

Creating a new Crossplane provider using Upjet is a streamlined process that allows you to extend Crossplane’s capabilities with minimal setup. Follow the official Upjet documentation to get started.

In our demonstration at the KCD Austria tech talk, we showcased how to build a Dynatrace provider. Below are the detailed steps we followed.

Step 1: Adjust the Makefile

The Makefile needs to reference the official Dynatrace Terraform module. This adjustment allows Upjet to use the correct source when generating the provider.

Here’s an example configuration:


export TERRAFORM_PROVIDER_SOURCE ?= dynatrace-oss/dynatrace
export TERRAFORM_PROVIDER_REPO ?= https://github.com/dynatrace-oss/terraform-provider-dynatrace
export TERRAFORM_PROVIDER_VERSION ?= 1.66.0
export TERRAFORM_PROVIDER_DOWNLOAD_NAME ?= terraform-provider-dynatrace
export TERRAFORM_PROVIDER_DOWNLOAD_URL_PREFIX ?= https://releases.hashicorp.com/$(TERRAFORM_PROVIDER_DOWNLOAD_NAME)/$(TERRAFORM_PROVIDER_VERSION)
export TERRAFORM_NATIVE_PROVIDER_BINARY ?= terraform-provider-dynatrace_v1.66.0
export TERRAFORM_DOCS_PATH ?= docs/resources

These settings specify the source, version, and download paths for the Dynatrace Terraform provider that Upjet will wrap as a Crossplane provider.

Step 2: Configure Provider Resources

Set Up the Provider Config

To configure the connection details, we need to modify internal/clients/dynatrace.go to reference the secret structure expected for the provider. In this case, define tenantURL and apiToken for Dynatrace connectivity:


const (
   tenantURL = "dt_env_url"
   apiToken  = "dt_api_token"
)

Then, reference these credentials in the TerraformSetupBuilder:


// TerraformSetupBuilder builds a Terraform setup function, returning provider configuration.
func TerraformSetupBuilder(version, providerSource, providerVersion string) terraform.SetupFn {
    return func(ctx context.Context, client client.Client, mg resource.Managed) (terraform.Setup, error) {
        ...
        // Set credentials in the provider configuration.
        ps.Configuration = map[string]any{}
        if v, ok := creds[tenantURL]; ok {
            ps.Configuration[tenantURL] = v
        }
        if v, ok := creds[apiToken]; ok {
            ps.Configuration[apiToken] = v
        }
    }
}

Define External Name Configurations

To identify external names for resources, update config/external_name.go by adding mappings for the Dynatrace resources:


// ExternalNameConfigs contains all external name configurations for this provider.
var ExternalNameConfigs = map[string]config.ExternalName{
    "dynatrace_alerting":           config.IdentifierFromProvider,
    "dynatrace_email_notification": config.IdentifierFromProvider,
    "dynatrace_json_dashboard":     config.IdentifierFromProvider,
    "dynatrace_metric_events":      config.IdentifierFromProvider,
}

This setup ensures that each resource is correctly identified using the provider’s unique identifier.

Add Custom Configurations for Resources

For each resource, create a corresponding config subfolder and add a config.go file with a Configure function. This function customizes the resource’s configuration and short group name as needed:

config/alerting/config.go


package alerting
import "github.com/crossplane/upjet/pkg/config"
// Configure customizes individual resources.
func Configure(p *config.Provider) {
  p.AddResourceConfigurator("dynatrace_alerting", func(r *config.Resource) {
    r.ShortGroup = "alerting"
  })
}

Repeat this for other resources, such as Dashboard, Event, and Notification.

config/dashboard/config.go


package dashboard
import "github.com/crossplane/upjet/pkg/config"
// Configure customizes individual resources.
func Configure(p *config.Provider) {
  p.AddResourceConfigurator("dynatrace_json_dashboard", func(r *config.Resource) {
    r.ShortGroup = "dashboard"
  })
}

config/event/config.go


package event
import "github.com/crossplane/upjet/pkg/config"
// Configure customizes individual resources.
func Configure(p *config.Provider) {
  p.AddResourceConfigurator("dynatrace_metric_events", func(r *config.Resource) {
    r.ShortGroup = "event"
  })
}

config/notification/config.go


package notification
import "github.com/crossplane/upjet/pkg/config"
// Configure customizes individual resources.
func Configure(p *config.Provider) {
  p.AddResourceConfigurator("dynatrace_email_notification", func(r *config.Resource) {
    r.ShortGroup = "notification"
  })
}

Register Custom Configurations

To ensure these custom configurations are applied, register each Configure function in config/provider.go:


import (
    "github.com/xoanmi/provider-dynatrace/config/alerting"
    "github.com/xoanmi/provider-dynatrace/config/event"
    "github.com/xoanmi/provider-dynatrace/config/dashboard"
    "github.com/xoanmi/provider-dynatrace/config/notification"
)
for _, configure := range []func(provider *ujconfig.Provider){
    alerting.Configure,
    event.Configure,
    notification.Configure,
    dashboard.Configure,
} {
    configure(pc)
}

This setup allows Upjet to apply each resource’s configuration during the provider generation process, customizing each resource group as specified.

Step 3: Generate the Code

Once all the necessary configurations are in place, you’re ready to generate the provider code by running the following command:

The make generate command will use the settings specified in the previous steps to:

  • Generate the Crossplane provider code based on the Terraform provider configurations.
  • Create the necessary Kubernetes Custom Resource Definitions (CRDs) for each resource, allowing Crossplane to manage them.

Running make generate will produce output similar to the following, showing the installation of required tools and the generation of the provider schema and resource CRDs:


➜ make generate
11:38:21 [ .. ] installing terraform darwin-arm64
…
11:38:22 [ OK ] installing terraform darwin-arm64
11:38:22 [ .. ] generating provider schema for dynatrace-oss/dynatrace 1.66.0
11:38:24 [ OK ] generating provider schema for dynatrace-oss/dynatrace 1.66.0
11:38:26 [ .. ] go generate linux_arm64

Generated 4 resources!
11:38:59 [ OK ] go generate linux_arm64
11:38:59 [ .. ] go mod tidy
11:39:00 [ OK ] go mod tidy

➜ tree package/crds
package/crds
├── alerting.crossplane.io_alertings.yaml
├── dashboard.crossplane.io_dashboards.yaml
├── dynatrace.crossplane.io_providerconfigs.yaml
├── dynatrace.crossplane.io_providerconfigusages.yaml
├── dynatrace.crossplane.io_storeconfigs.yaml
├── event.crossplane.io_events.yaml
└── notification.crossplane.io_notifications.yaml

With the generated code and CRDs in place, you can now deploy the provider and begin managing Dynatrace resources in your Kubernetes environment.

Step 4: Deploy and run

In the setup, we’re going to run the operator locally while applying and connecting to a Kubernetes cluster (where this cluster runs doesn’t matter, as long as it’s reachable).

First, we apply the CRDs make generate has created.


kubectl apply -f package/crds

Once this is done, you can run the operator itself.


make run

The last missing part is adding the credentials the operator needs to connect to the chosen Dynatrace tenant. The necessary token is in the Access management documentation.

Create the namespace and a secret containing your brand-new access token.


apiVersion: v1
kind: Secret
metadata:
  name: example-creds
  namespace: crossplane-system
type: Opaque
stringData:
  credentials: |
    {
     "dt_env_url": "https://my-tenant.com",
     "dt_api_token": "my-secret-token"
    }

kubectl create namespace crossplane-system
kubectl apply -f example-creds.yaml

Finally, the last puzzle piece, the ProviderConfig, can be created:


apiVersion: dynatrace.crossplane.io/v1beta1
kind: ProviderConfig
metadata:
  name: default
  namespace: crossplane-system
spec:
  credentials:
    source: Secret
    secretRef:
      name: example-creds
      namespace: crossplane-system
      key: credentials

All done! The previously generated CRDs are now available and their objects in your Kubernetes cluster will result in entities and changes in your Dynatrace tenant. Try it out with a dashboard resource!


apiVersion: dashboard.crossplane.io/v1alpha1
kind: Dashboard
metadata:
  name: example-dashboard
  namespace: crossplane-system
spec:
  forProvider:
    contents: |
      {
        "dashboardMetadata": {
          "name": "Our small example dashboard",
          "owner": "my@mail.com",
          "preset": true,
          "hasConsistentColors": true
      },
      "tiles": [
        {
          More config…
        }
     }

kubectl apply -f example-dashboard.yaml

FAQ

Several key questions were raised during the talk and in the following discussions. We’ve summarized the main points:

Q: Why use Crossplane when Terraform can do the same thing?

A: It’s not a matter of “should” versus “shouldn’t.” If you’re already invested in Kubernetes, Crossplane allows you to manage cloud resources while leveraging the same tooling you use to deploy, maintain, and monitor your applications. This makes Crossplane highly convenient for teams already embedded in the Kubernetes ecosystem.

Additionally, these tools don’t exclude each other. One is used to build platforms, and the other is a command-line tool. Their potential use cases differ quite a lot.

Q: How is state management handled?

A: With the Upjet approach, you’re essentially bridging two worlds. Kubernetes manages the state of each object through its etcd system. Simultaneously, the Crossplane operator runs Terraform in the background, continuously reconciling the state between the Kubernetes Custom Resource (CR) and the Terraform-managed infrastructure.

Q: Can I use the Upjet approach in production?

A: Yes, you can, but remember that the provider uses Terraform under the hood. This means that during each reconciliation loop, a terraform plan and terraform apply run. Due to the nature of these continuous operations, managing a large number of resources this way could demand significant resources.

While the Upject project is very good at translating the provider, some things need to be added manually. The concept of Kubernetes labels simply doesn’t exist in Terraform. If you want to utilize them, you need to implement them yourself.

Q: How does the mapping between Terraform objects and Kubernetes Custom Resources (CRs) work?

A: The mapping is defined in the `/config` folder when configuring the provider. Here, we specify the relationship between the Terraform object and the corresponding Kubernetes CR. Running the `make generate` command triggers the generation of all necessary code, including the API, client, provider, and CRD (Custom Resource Definition). This allows Kubernetes to manage the Terraform-defined resource seamlessly.

Q: I read about the proposal for Crossplane v2.0. Do you know if that will impact the described provider creation process?

A: Recently, the Crossplane developers created a draft for the next version of Crossplane. Here, they talk in-depth about how they want to change composite resources and their structure. This will impact provider creation since they reconcile aforementioned resources. This section discusses the proposed changes. The developers also plan on keeping things backward compatible. For now, we
have to wait and see what the final implementation looks like.

Get started

Crossplane enables a seamless cloud-native approach for managing any cloud resource by extending the Kubernetes API. By leveraging Kubernetes as a control plane and using Crossplane compositions, you can declaratively define and automate your entire observability stack.

It’s easy to get started, all you need to start is

If you’re interested in diving deeper, you can check out the following resources from our session:

We hope this talk inspired you to explore Crossplane for your infrastructure automation needs and provided valuable insights into building observability solutions using the power of Kubernetes and declarative infrastructure.

The post Observability as Code: DIY with Crossplane appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/observability-as-code-diy-with-crossplane/feed/ 0
Dynatrace on Microsoft Azure in Australia enables regional customers to leverage AI-powered observability https://www.dynatrace.com/news/blog/dynatrace-saas-in-australia-on-azure/ https://www.dynatrace.com/news/blog/dynatrace-saas-in-australia-on-azure/#respond Mon, 28 Oct 2024 22:14:35 +0000 https://www.dynatrace.com/news/?p=66157 Dynatrace | Azure

As modern multicloud environments become more distributed and complex, having real-time insights into applications and infrastructure while keeping data residency in local markets is crucial. Dynatrace on Microsoft Azure allows enterprises to streamline deployment, gain critical insights, and automate manual processes. The result? Optimized performance and enhanced customer experiences. As of October 2024, Dynatrace is […]

The post Dynatrace on Microsoft Azure in Australia enables regional customers to leverage AI-powered observability appeared first on Dynatrace news.

]]>
Dynatrace | Azure

As modern multicloud environments become more distributed and complex, having real-time insights into applications and infrastructure while keeping data residency in local markets is crucial. Dynatrace on Microsoft Azure allows enterprises to streamline deployment, gain critical insights, and automate manual processes. The result? Optimized performance and enhanced customer experiences.

As of October 2024, Dynatrace is available on Microsoft Azure Australia East region, enabling joint customers to maintain a local SaaS presence. For Dynatrace customers, this means their data and end users in the region will benefit from faster time to value and deeper integration with the Microsoft technology stack to help comply with local data privacy and security requirements.

The move to cloud and data residency in local markets

One of the most significant advantages of this launch is maintaining data residency in local markets. By keeping data within the region, Dynatrace ensures compliance with data privacy regulations and offers peace of mind to its customers. This local SaaS presence minimizes latency and maximizes the speed and reliability of data access.

As a SaaS vendor, Dynatrace carefully manages its deployments across different regions, assuring the efficient and optimal use of infrastructure to serve and support Dynatrace platform customers. The regional reach of the Dynatrace AI-powered platform as a SaaS on Microsoft Azure has now expanded to Australia.

Transforming enterprise operations with Azure Native Dynatrace Service

With the availability of the Dynatrace platform on Microsoft Azure in Australia, regional customers are also able to take advantage of the integration of Azure Native Dynatrace Service.

This integration provides customers with a streamlined and automated setup and configuration of Dynatrace through the Azure marketplace/portal. Additionally, customers are able to maximize their Microsoft investment leveraging Microsoft Azure Consumption Commitment (MACC) to purchase Dynatrace directly through the Azure marketplace.

“Our partner Dynatrace works alongside Microsoft to provide deep insights and automation across the technology stack to enhance operational efficiency for organizations. The Azure Native Dynatrace Service, available in Azure Marketplace, uses Microsoft Azure AI and Dynatrace AI observability to ensure enterprise applications and services are high-performing, reliable, and secure.”

– Yvonne Muench, Sr. Director - Marketplace & ISV Journey, Microsoft

Maximize your cloud estate with AI-powered observability and security

If you’re an existing Dynatrace customer, please contact us to learn how to upgrade to Azure Native Dynatrace Service. An overview of how to upgrade is also available in our guide, Upgrade to Azure Native Dynatrace Service. Or, contact our team to discuss your specific use case. Dynatrace offers a comprehensive set of services to support your migration to the cloud.

Get started with the AI-powered Dynatrace platform on Microsoft Azure in Australia

Interested in exploring the benefits of Azure Native Dynatrace Service? Start by visiting the Azure Marketplace or contact us for more information.

The post Dynatrace on Microsoft Azure in Australia enables regional customers to leverage AI-powered observability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-saas-in-australia-on-azure/feed/ 0
Build systems more reliably with Dynatrace: Chaos Engineering https://www.dynatrace.com/news/blog/build-systems-more-reliably-with-dynatrace-chaos-engineering/ https://www.dynatrace.com/news/blog/build-systems-more-reliably-with-dynatrace-chaos-engineering/#respond Wed, 21 Aug 2024 15:29:30 +0000 https://www.dynatrace.com/news/?p=65273 Dynatrace introduces support for OpenTelemetry histograms

The previous blog post in this series discussed the benefits of implementing early observability and orchestration of the CI/CD pipeline using Dynatrace. This approach enhances key DORA metrics and enables early detection of failures in the release process, allowing SREs more time for innovation. This blog post explores the Reliability metric, which measures modern operational […]

The post Build systems more reliably with Dynatrace: Chaos Engineering appeared first on Dynatrace news.

]]>
Dynatrace introduces support for OpenTelemetry histograms

The previous blog post in this series discussed the benefits of implementing early observability and orchestration of the CI/CD pipeline using Dynatrace. This approach enhances key DORA metrics and enables early detection of failures in the release process, allowing SREs more time for innovation. This blog post explores the Reliability metric, which measures modern operational practices.

Why reliability?

DORA introduced the Reliability metric as “the fifth DORA metric” to address the fact that many organizations improved release times and deployment frequency without adequately considering the reliability of new releases. These releases often assumed ideal conditions such as zero latency, infinite bandwidth, and no network loss, as highlighted in Peter Deutsch’s eight fallacies of distributed systems. To enhance reliability, testing the software under these conditions is crucial to prepare for potential issues by leveraging chaos engineering or similar tools. Chaos engineering is a practice that extends beyond traditional failure testing by identifying unpredictable issues. While it is powerful, it presents several challenges that affect its adoption. In this blog post, we delve into these challenges and explore how Dynatrace can address them to enhance the reliability of released software.

Challenges

Dynatrace Observability to improve Reliability slide

Limited awareness of cross-service interactions

In chaos engineering, “hypothesis” denotes the expected behavior of an application under defined conditions or stressors, whether within their owned service or across services. It forms the cornerstone of chaos engineering experiments. In distributed or microservices environments, application teams often lack visibility into how their service will perform under diverse conditions across other services or the entire system. This limitation affects the quality of hypotheses, leading to frequent revisions or inaccurate assumptions.

“It’s not about chaos – it’s about reliability” – Andrus Kolton, CTO & founder of Gremlin

Dynatrace addresses this challenge by offering comprehensive observability across the entire environment. With Dynatrace, teams can seamlessly monitor the entire system, including network switches, database storage, and third-party dependencies. This visibility can be cross-checked in real time using features like Smartscape topology or Service Flow. By leveraging these resources, SREs can formulate more informed hypotheses based on the behavior of their service or application.

Insights: Unclear starting system state

Along with hypotheses, it is imperative to have a clear starting state to validate the experiments’ success. These are referred to as baseline metrics for the starting state, and without appropriate insights into the application ecosystem, it is difficult to form hypotheses. Such baselines constitute a few metrics like:

  1. What are the top five problems in your application – CPU spikes, slow response, database connections bottleneck, etc.
  2. The problems that take maximum time to resolve – lowest MTTR.
  3. Impact of fewer resources, for example, CPU and disk, available to different services and applications.
  4. More critical services that are likely to take other services down.
With Dynatrace, you can use DQL to identify insights from your application's behavior in production and determine the starting system state.
With Dynatrace, you can use DQL to identify insights from your application’s behavior in production and determine the starting system state.

For instance, the above dashboard offers visibility into service and distributed systems through different lenses, such as the slowest MTTR, top problems, or services frequently impacting others. This data simplifies establishing the starting state for chaos experiments or determining the priority services requiring reliability tests and the specific reliability standards to test against.

Blast radius and risk of prolonged outages

While tests in pre-production environments provide valuable insights, proper validation occurs when chaos experiments are conducted in the production environment. However, SREs are hesitant to use chaos engineering during testing or validation stages due to the risk of incorrect hypotheses, potentially impacting multiple services in the production system and leading to prolonged outages. Therefore, their primary concern is to avoid such outages and unnecessary damage to the production system, leading to low adoption of chaos engineering and hindering efforts to improve reliability.

These outages can be mitigated by limiting the scope of chaos engineering tests, known as the “blast radius,” and gradually expanding it as impacted entities are correctly identified and fixed. Dynatrace automatically discovers all components and dependencies within complex technology stacks end-to-end, identifying billions of causal dependencies across websites, applications, services, processes, hosts, networks, and infrastructure within minutes. With this comprehensive monitoring, SREs can utilize holistic monitoring, where situational awareness and Davis® AI-based alerting transition from correlation to causation-based analysis. This enables different teams to understand the semantics of cascading problems, or the domino effect, and pinpoint the root cause for actionable fixes.

In the screenshot below, a chaos engineering scenario introduced latency and resource stress on the “easytrade” demo application. This scenario introduces a 7-second delay for random requests and causes CPU stress on the container. This helps the app team observe the application’s behavior under these conditions and identify any bottlenecks.

Chaos Engineering with Dynatrace screenplay

In the video below, Davis AI identifies the interdependencies and pinpoints that database slowness impacted multiple services, identifying it as the root cause. Additionally, the details of the new deployment and chaos experiment were included in the root cause analysis. By leveraging service tools and visual resolution, developers can work towards building robust and reliable services. Similarly, by gradually increasing the blast radius, developers can enhance the reliability of the entire application.
Video thumbnail

Lastly, the SRE team can leverage Dynatrace workflows to automate outages, ensuring virtually no downtime for services or applications.

Conclusion

Ensuring builds are created accurately the first time is crucial, even if it means testing them by injecting failures or remote scenarios. This enhances build reliability and equips SREs and operations teams to swiftly resolve potential issues that could lead to outages. As John Allspaw said, “Incidents are unplanned investments; their costs have already been incurred. Your org’s challenge is to get ROI on those events.”

With Dynatrace advanced monitoring capabilities, SRE teams can confidently overcome challenges and embrace Chaos Engineering. This ensures optimal performance and reliability, enabling organizations to confidently navigate today’s dynamic IT landscapes.

What’s next

In this blog series’ final installment, we’ll explore how Davis AI root cause analysis can be used to autonomously manage your application’s day-to-day operations with the help of Dynatrace workflows.

The post Build systems more reliably with Dynatrace: Chaos Engineering appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/build-systems-more-reliably-with-dynatrace-chaos-engineering/feed/ 0
Three ways Dynatrace can help to drive innovation through cloud modernization https://www.dynatrace.com/news/blog/dynatrace-for-executives-cloud-modernization/ https://www.dynatrace.com/news/blog/dynatrace-for-executives-cloud-modernization/#respond Thu, 18 Jul 2024 13:30:14 +0000 https://www.dynatrace.com/news/?p=64730 Dynatrace for Executives: Cloud Modernization

As executives, we drive change, balancing modernization speed with its risks. Technology—both a blessing and a curse—not only propels businesses forward but also adds complexity as developers introduce new innovations to enhance customer services and competitiveness. Anticipate future customers’ needs Anticipating customer needs three to five years ahead helps to reduce wasted investments into “wants” […]

The post Three ways Dynatrace can help to drive innovation through cloud modernization appeared first on Dynatrace news.

]]>
Dynatrace for Executives: Cloud Modernization

As executives, we drive change, balancing modernization speed with its risks. Technology—both a blessing and a curse—not only propels businesses forward but also adds complexity as developers introduce new innovations to enhance customer services and competitiveness.

Anticipate future customers’ needs

Anticipating customer needs three to five years ahead helps to reduce wasted investments into “wants” and directs them toward “needs” that future-proof the business.

This mentality has driven me to continuously innovate and reinvent Dynatrace®. My ongoing evaluation of how technology changes the way digital services are architected allowed me to recognize early on that change is on the horizon. The rise of cloud-native technologies, the convergence of observability and security, and the demand for actionable insights required a new approach to managing data at an exabyte scale, as existing databases could no longer keep up.

Change is constant

In our fast-paced world, success requires thinking big but acting small to create value quickly and sustainably. For cloud modernization, this means executives must change how software is built, operated, and secured; improve collaboration processes; and increase automation.

Dynatrace gives executives an indispensable platform for driving this change in the following three ways:

  • Enabling a modern AIOps strategy,
  • Accelerating software delivery, and
  • Making scarce engineering resources more productive.
Key insights for executives
  • Modern AIOps and AISecOps from Dynatrace get us closer to NoOps and NoSoc than ever with help of hypermodal AI
  • Early investment into automation pays off, and the 100 ready-made use cases  from Dynatrace accelerate software delivery with confidence
  • Extend to the left has become the modern shift left, and Dynatrace accelerates productivity with contextual analytics, AI, automation, and platform engineering

1. Go beyond traditional AIOps

The first wave of AIOps investment was about “noise reduction.” This has been helpful but falls short of the potential offered by the preventive NoOps and NoSOC approaches that many executives seek AI to enable.

With current hype causing a resurrection in AI investment, it is tempting to believe that this time, machine learning and generative AI will fulfill the promises of the past. However, while the advances in machine learning-based AI are a huge step up for many use cases, it is still problematic to apply it to prevent incidents and errors in IT systems. Why? Because training an AI requires errors, failures, and behaviors to occur many times to ‘learn’. While the exact numbers may have been reduced by the advances in generative AI, which executive wants to have service outages just to train AI to prevent them in the future? Even if it was possible to arrive at a trained model, it would quickly become obsolete as services get updated and new features introduced.

As we consider a way forward, I urge all executives to recognize that we are in the trough of disillusionment in the AI hype cycle. This is good news, as it allows us to think more rationally. We need to understand that there are multiple types of AI, each suited for different purposes.

Dynatrace is uniquely designed to help executives elevate their AIOps – and AISecOps strategy – to a different level by combining multiple types of AI in a single framework known as hypermodal AI: the power of predictive AI, causal AI, and generative AI for observability, security, and business use cases. Proven by thousands of customers in large-scale IT deployments, this approach delivers greater speed, automation, and precision.

Our hypermodal AI automatically infers the root cause of issues based on a real-time updated graph without needing to learn. Now, it is more feasible than ever to automate workflows for self-healing, security investigation, and preventive operations to deliver great software with confidence, all while enhancing security measures and boosting productivity.

2. Accelerate software delivery

One of the best features of the cloud and Kubernetes® is achieving most availability needs with minimal effort, a major improvement over the classic datacenter model. This allows executives to focus on accelerating software delivery. However, the inverse Pareto principle applies: achieving the final 20% of flawless, secure services requires 80% of the effort.

That’s why APIs have become my favorite feature of the cloud as the key to automate and orchestrate. This is where Dynatrace comes in. Dynatrace integrates with the cloud ecosystem and DevOps toolchain to enhance automation across software delivery, resilience, and security throughout the software lifecycle.<

Throughout the ten years since we embraced NoOps at Dynatrace, I understood the temptation to favor releasing new features over investing in automation. Automation always paid off. We have since developed over 100 ready-made use cases to support platform engineering across the software delivery lifecycle. From development and release to operation and flaw prevention, prediction, and resolution, Dynatrace offers a robust data analytics-driven automation platform.

We’ve seen the many benefits of investing in automation, including the following capabilities:

  • Releasing faster and securely with automated quality and security gates
  • Catching bugs earlier, before customers experience them
  • Preventing issues with predictive operations
  • Avoiding unnecessary high consumption and cost with causal and predictive auto-scaling
  • Empowering developers with context-rich insights derived from self-service observability and security
  • Orchestrating more intelligently with real-time user behavior and business data

In a nutshell, Dynatrace allows executives to accelerate software delivery with confidence.

Dynatrace Dashboards: visualize your complex hybrid cloud environments in real time, gaining insights into security and business performance.

3. Increase teams’ productivity

As Dynatrace CTO, one of the questions constantly on my mind is: how can I enable my team to be more productive?

Over the past 15 years, most of us have embraced the “shift left” ethos to empower software developers. The earliest iteration of this was the “you build it, you run it” mentality. However, given the responsibilities of creating enterprise-scale and secure software, the “extend left” ethos proves to be more successful and fitting for cloud modernization.

Extend to the left: The modern “shift left”

“Extend left” refers to sharing responsibility amongst developers and operations teams, through adding more self-service for developers while retaining consistency, tooling and knowledge management with central teams.

As neither full decentralization nor full centralization will be effective, a hybrid model, supported by platform engineering approaches, is much more likely to succeed. Centralizing the necessary expert knowledge within a platform engineering team enables rapid, secure, and safe software delivery. At the same time, this approach decentralizes innovation, making it accessible to many.

Dynatrace was created to enable precisely this approach, leveling up developer experience by providing self-service capabilities while allowing central safety and oversight maintenance. This gives executives the best of both worlds: decentralized autonomy supported by centralized governance and control.

Armed with the use cases across the three areas outlined here, executives can modernize their cloud operations faster and equip their teams with the capabilities they need to accelerate innovation confidently. As a result, they will be better placed to anticipate change and continuously reinvent their organization to stay ahead of the market.

Follow along the new “Dynatrace for Executives” blog series. In the coming weeks, I’ll dive deeper into each of the nine executive use case areas to help you unlock the potential of Dynatrace.
Want to learn more about all nine use cases? See the overview on the homepage.

The post Three ways Dynatrace can help to drive innovation through cloud modernization appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-for-executives-cloud-modernization/feed/ 0
Business observability: From IT monitoring to driving digital transformation https://www.dynatrace.com/news/blog/business-observability-drives-digital-transformation/ https://www.dynatrace.com/news/blog/business-observability-drives-digital-transformation/#respond Thu, 11 Jul 2024 16:24:23 +0000 https://www.dynatrace.com/news/?p=64682 Dynatrace AI-powered observability is now on Google Cloud

As organizations adopt more cloud-native technologies, traditional IT monitoring is no longer up to the task of supporting wider business needs. Organizations need to shift toward more sophisticated models of monitoring and managing IT operations. The best way to accomplish this upgrade is to implement a business observability strategy.

The post Business observability: From IT monitoring to driving digital transformation appeared first on Dynatrace news.

]]>
Dynatrace AI-powered observability is now on Google Cloud

Cloud-native technologies are driving the need for organizations to adopt a more sophisticated IT monitoring approach to satisfy the competitive demands of modern business. Business observability is emerging as the answer.

The ongoing drive for digital transformation has led to a dramatic shift in the role of IT departments. They’ve gone from just maintaining their organization’s hardware and software to becoming an essential function for meeting strategic business objectives. Today, IT services have a direct impact on almost every key business performance indicator, from revenue and conversions to customer satisfaction and operational efficiency.

As a result, organizations have been forced to reevaluate what success looks like for the modern IT department and how they monitor and manage the performance of IT services.

Seeking insights from data

Every organization depends on data to make decisions. However, too often teams are forced to rely on disjointed data that lacks context, which leads to poor decision-making and wasted resources.

This problem has worsened as enterprise operational complexity has grown. In today’s digital-first world, data resides across dozens of different IT systems, from critical business applications to the modern cloud platforms that underpin them. Connecting the dots between these various silos of data to understand the relationship between the health of IT services and the business outcomes they enable has become a particular challenge.

The journey toward business observability

Traditional IT monitoring that relies on a multitude of tools to collect, index, and correlate logs from IT infrastructure, networks, applications, and security systems is no longer effective at supporting the need of the wider organization for business insights. This traditional approach presents key performance metrics in an isolated and static way, providing little or no insight into the business impact or progress toward the goals systems support. Often, these metrics are unable to even identify trends from past to present, never mind helping teams to predict future trends.

As a result, organizations need to shift toward more sophisticated models of monitoring and managing IT operations. With hybrid and multi-cloud architectures rendering organizations’ environments more complex and distributed, cloud observability has become increasingly important. Likewise, integrating metrics and traces with log data helps to identify crucial context that reveals the interconnections among and importance of signals from all levels of the network. These capabilities are essential to providing real-time oversight of the infrastructure and applications that support modern business processes. With cloud observability, organizations can make data-driven decisions to improve the health of their IT services and proactively mitigate potential risks.

Partners such as Deloitte provide key expertise in cloud observability and are instrumental for many organizations embarking on digital transformation. By leveraging Deloitte’s strategic insights, businesses can align their IT investments more closely with their overarching business objectives, driving both efficiency and growth.

However, the journey doesn’t end there. The final stage is developing true business observability. Business observability ensures that all IT activity and investment is aligned with an organization’s strategic business objectives by enabling superior data-driven decision-making. It provides insights to help address not only operational issues such as cost reduction and risk mitigation but also customer-centric issues such as optimizing user journeys and creating personalized experiences.

Five ways business observability drives impact

There are several key advantages to making the transition to business observability, from mitigating the risks of adopting new cloud architectures and the challenges of data sovereignty, to rightsizing the IT estate and implementing greener technology. Five of the most important benefits of modern business observability are identified below.

  1. Operational optimization. Across all sectors, system performance, infrastructure reliability, and transaction speeds are essential, whether it’s grid management for the energy industry or supply chain integration for retailers. An effective business observability strategy can help to meet these requirements by stabilizing the entire application stack, reducing operational expenditure, and preventing downtime in critical business systems.
  2. Security and compliance. Organizations need to continually track transactions, user behaviors, and alerts to maintain security and compliance by detecting anomalies and indicators of potentially fraudulent activity or breaches. Business observability provides a cohesive approach to meeting these goals by offering a complete end-to-end view of application threats and vulnerabilities, assessed according to the level of potential risk to the enterprise.
  3. Optimized experiences. By integrating data on user interactions, omnichannel behaviors, customer journeys, and purchasing patterns, organizations can take more effective action to deliver consistent and personalized experiences.
  4. Resource optimization. Tracking and analyzing data from multiple sources enables businesses to optimize the performance of their IT services and prevent unnecessary downtime, while also reducing unnecessary resource consumption and carbon emissions through efficient asset utilization.
  5. Agility and innovation. Organizations need oversight of the entire innovation pipeline, from ideation to implementation, to identify bottlenecks and streamline development and testing processes. Mature business observability capabilities allow businesses to reduce time-to-market for innovation by streamlining the product development cycle, while also providing key insights into user needs and the feasibility of potential feature additions.

Ultimately, organizations with mature business observability capabilities are better placed to use IT as a catalyst to drive better outcomes, streamline their operations, and mitigate risks, while unlocking greater customer satisfaction. This will help to place them at the forefront of the digital transformation landscape.

For further insights on the practical steps organizations can take to progress along their own journey from IT monitoring to business observability, download the full whitepaper.

The post Business observability: From IT monitoring to driving digital transformation appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/business-observability-drives-digital-transformation/feed/ 0
The 3 biggest Kubernetes deployment mistakes you can make https://www.dynatrace.com/news/blog/the-3-biggest-kubernetes-deployment-mistakes/ https://www.dynatrace.com/news/blog/the-3-biggest-kubernetes-deployment-mistakes/#respond Tue, 02 Apr 2024 14:55:42 +0000 https://www.dynatrace.com/news/?p=63292 Kubernetes deployment graphic

The last decade brought a wave of digital transformation that accelerated the move towards cloud-native tech, specifically Kubernetes. We could talk about this move’s benefits (and downsides!) for more than a few blogs, but that’s not why we’re here. We’re here to talk about the less savory side of cloud-native transformation—you know, all the things […]

The post The 3 biggest Kubernetes deployment mistakes you can make appeared first on Dynatrace news.

]]>
Kubernetes deployment graphic

The last decade brought a wave of digital transformation that accelerated the move towards cloud-native tech, specifically Kubernetes. We could talk about this move’s benefits (and downsides!) for more than a few blogs, but that’s not why we’re here. We’re here to talk about the less savory side of cloud-native transformation—you know, all the things that can go wrong. More specifically, all the Kubernetes deployment mistakes you can make when moving to Kubernetes—and how to avoid making them in the first place.

As someone who has worked deep in the coding trenches with developers my whole life, I’ve hand-picked the top three mistakes you can make when moving to Kubernetes. So, without further ado, let me share these hard-earned mistakes you should avoid like the plague when moving to Kubernetes!

Kubernetes deployment mistake #1: Managing Kubernetes from the command line

Kubernetes deployments almost feel like magic the first time you get them working. You use a (hopefully) short YAML file to specify the application you want to run, and Kubernetes just makes it so. Make a change to the file, apply it, and it will update in near real-time.

But as powerful as kubectl is, and as instructive as it can be to explore Kubernetes using it, you should not come to rely on kubectl too much. Of course, you’ll return to it (or its amazing cousin, k9s) when you need to troubleshoot issues in Kubernetes, but don’t use it to manage your cluster.

Kubernetes was made for the configuration-as-code paradigm, and all those YAML files belong in a Git repo. You should commit any and all of your desired changes to a repo and have an automated pipeline deploy the changes to production. Some of your options include:

Kubernetes deployment mistake #2: Forgetting all about resources

Let’s assume all your workloads are up and running with all the goodness of Kubernetes and configuration as code. But now you’re orchestrating containers, not virtual machines. How do you ensure they get the CPU and RAM they need? Through resource allocation!

Resource requests

What happens if you forget to set resource requests?

Kubernetes will pack all your pods (“workloads” in Kubernetes-speak) into a handful of nodes. They won’t get the resources they need. The cluster won’t scale itself up as needed.

What are resource requests?

Resource requests tell the scheduler how many resources you expect your application to consume. When assigning pods to nodes, Kubernetes budgets them so that the node’s resources meet all of their requirements.

Resource limits

What happens if you forget to set resource limits?

A single pod may consume all the CPU or memory available on the node, causing its neighbors to be starved of CPU or hit Out of Memory errors.

What are resource limits?

Resource limits let the container runtime know how many resources you allow your application to consume. For the CPU limit, your application will be able to get that much CPU time but no more. Unfortunately (for the application), if it hits the memory limit, it will be OOMKilled by the container runtime.

So, go ahead and define requests and limits for each of your containers. If you aren’t sure, just take a guess, and keep in mind that the safe side is higher. Whether you’re certain or not, make sure to monitor actual resource usage by your pods and containers by using your cloud provider or APM tools.

Kubernetes deployment mistake #3: Leaving the developers behind

Immutable infrastructure and clean upgrades. Easy scalability. Highly available, self-healing services. Kubernetes provides you with lots of value directly out of the box. Unfortunately, this value might not be a priority for the developers working on your product. Your developers have other concerns:

  • How do I build and run my code?
  • How do I understand what my code is doing in development, testing, and integration?
  • How do I investigate bugs reported in QA and production environments?

For many of these tasks, Kubernetes pulls the rug out from under the developer. Running development environments locally is much harder because many dev and test workloads are moved to the cloud. The code-level visibility developers rely on is often poor in these environments, and direct access to the application and its filesystem is virtually impossible.

Successful Kubernetes adoption requires the right tools

To lead a successful adoption of a new platform such as Kubernetes, you need everyone to see the value in it. But don’t forget that developers require the right tools to keep up with their code and understand what it’s doing as it’s running.

Get started on your Kubernetes journey with Dynatrace.

The post The 3 biggest Kubernetes deployment mistakes you can make appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/the-3-biggest-kubernetes-deployment-mistakes/feed/ 0
Dynatrace OTel Collector distribution amplifies OpenTelemetry integration for scalable, production-ready observability https://www.dynatrace.com/news/blog/dynatrace-opentelemetry-collector-for-production-ready-observability/ https://www.dynatrace.com/news/blog/dynatrace-opentelemetry-collector-for-production-ready-observability/#respond Mon, 18 Mar 2024 16:15:54 +0000 https://www.dynatrace.com/news/?p=63116 OpenTelemetry and Dynatrace make a winning combination

Dynatrace is announcing the Dynatrace OTel Collector distribution, an OpenTelemetry Collector with components that are fully tested, reliable, secure, and supported for use in production and enterprise-scale environments. OpenTelemetry provides a flexible, standardized way to collect telemetry data, and its adoption is growing. The Dynatrace OTel Collector distribution hardens OpenTelemetry functions for organizations that need to maintain high-reliability service-level agreements. The Dynatrace OTel Collector is free and available to everyone. Dynatrace donates bug fixes and enhancements back to the OpenTelemetry project and provides full support for customers.

The post Dynatrace OTel Collector distribution amplifies OpenTelemetry integration for scalable, production-ready observability appeared first on Dynatrace news.

]]>
OpenTelemetry and Dynatrace make a winning combination

OpenTelemetry standardizes how organizations instrument, generate, and collect telemetry data for analysis and provides community-based support. Because of its flexibility, this open source approach to instrumenting and collecting telemetry data is becoming increasingly important in large-size organizations. But rigorous requirements for security, production readiness, scalability, and reliability can make adopting OpenTelemetry challenging for teams to maintain at enterprise scale.

To meet these requirements, Dynatrace has released its OpenTelemetry Collector distribution (Dynatrace OTel Collector), providing users with a fully tested, secure, and reliable OpenTelemetry collector for production and enterprise-scale environments. The distribution is fully supported for Dynatrace customers at standard and premium levels.

Growing OpenTelemetry adoption drives the need for production-level reliability, security, and stability

OpenTelemetry provides a standardized way for organizations to instrument, generate, and collect telemetry data. Organizations use it to collect and send data to a backend, such as Dynatrace, that can analyze software performance and behavior. As a result of its standardized, flexible approach, OpenTelemetry is growing fast in popularity.

In fact, according to technology intelligence analyst firm HG Insights, OpenTelemetry is the Cloud Native Computing Foundation’s second most active project after Kubernetes and is used by more than 1,600 organizations worldwide, with customer adoption growing at a compound annual growth rate of 59%. More than 60 vendors support OpenTelemetry.

Likewise, the OpenTelemetry Project Journey Report states that more than 9,000 contributors and 1,100 companies have contributed to OpenTelemetry since the project’s inception in 2019. Dynatrace is among the project’s top contributors.

Before OpenTelemetry and the W3C Trace Context open standard that underpins it, observability vendors had to reverse-engineer tracing libraries. Now, developers can build software libraries and use OpenTelemetry to add tracing and telemetry directly into them so an observability analytics backend, such as Dynatrace, can consume the data immediately.

Depending on their requirements, Dynatrace customers can ingest enriched telemetry data using OneAgent® or the Dynatrace OTel Collector to export to Dynatrace. OneAgent provides the enriched telemetry automatically. OpenTelemetry is natively supported by middleware and cloud providers and can ingest telemetry data directly. Used together, Dynatrace extends OpenTelemetry observability, and OpenTelemetry extends Dynatrace observability.

“At Dynatrace, we firmly believe that a thriving observability ecosystem benefits not only our customers but also our valued partners,” said Alois Reitbauer, Chief Technology Strategist at Dynatrace. “Our unwavering commitment to open source initiatives underscores our mission to foster transparency, collaboration, and innovation. Together, we shape the future of observability, empowering organizations to unlock actionable insights and elevate their cloud-native experiences.”

The Dynatrace OTel Collector distribution brings OpenTelemetry production readiness to everyone

Building on the flexibility and strengths of OpenTelemetry, the Dynatrace OTel Collector provides production-readiness for organizations that need to maintain highly reliable software environments and service-level agreements. Dynatrace does the work of maintaining the distribution, fixing bugs, and verifying that components, such as receivers and exporters, are stable, secure, and work reliably, saving developers time and resources. In addition, as a top contributor to the OpenTelemetry standard, Dynatrace donates bug fixes and enhancements back to the project.

Releasing the Dynatrace OTel Collector reinforces the company’s commitment to open source software and democratizing how organizations collect data from cloud infrastructure and applications. In open source tradition, the Dynatrace OTel Collector distribution is free and available for everyone.

“OpenTelemetry is gaining increasing importance in collecting telemetry data,” Reitbauer said. “With our collector distribution, we provide an enterprise-ready way for users to collect data reliably and leverage the power of the Dynatrace platform on their data.”

Using the Dynatrace OTel Collector, developers can safely and reliably collect telemetry data from different sources and send it to Dynatrace for analysis using the Grail™ data lakehouse and Davis® hypermodal AI. The Dynatrace OTel Collector also seamlessly integrates with Dynatrace OpenPipeline™, a stream-processing technology, to filter, process, and contextualize data for optimal use with Dynatrace if customers choose to use it as their backend.

As developers add OpenTelemetry to their applications, they can use the Dynatrace distribution to provide the instrumentation without needing to configure infrastructure. With deep analytics of traces from Dynatrace, developers have data in full context, which helps them easily debug instrumentation issues.

Dynatrace OTel Collector benefits free for everyone—full support for customers

The Dynatrace distribution of the OpenTelemetry Collector provides stable and reliable production-ready OpenTelemetry data collection for Dynatrace customers and noncustomers alike, for free. Dynatrace customers receive full support at standard and premium levels.

  • Dynatrace-verified collector components. Dynatrace verifies all the integrated collector components including receivers and exporters to ensure they operate seamlessly in production environments.
  • Innovation from the open source community. The OpenTelemetry Collector has an active community that constantly brings new features and innovation. The Dynatrace distribution introduces the latest features, but only after the Dynatrace team has fully tested them and can recommend them for production.
  • Sample configurations and documentation. Dynatrace Documentation includes example configurations for the key use cases, as well as for non-Dynatrace customers.
  • Immediate Security patches. Dynatrace contributes feature enhancements and fixes back to the community to benefit the entire open source project and all its users. However, with this distribution, Dynatrace can release critical patches faster and independently of the upstream OpenTelemetry release cycle.
  • Dynatrace standard or premium support. For Dynatrace customers, the Collector comes with proactive customer guidance and support, technical account management, flexible contact options, and proven expertise. With the help of support, customers can get an easy start with the Dynatrace OTel Collector.

Dynatrace OTel Collector distribution makes common use cases easier at enterprise scale

By ingesting OpenTelemetry data using a fully tested, secure, and reliable collector, teams can easily maintain service-level agreements. Following are a few common use cases that the Dynatrace OTel Collector distribution makes easier at enterprise scale.

  • Migrating logs from several sources. Organizations commonly use OpenTelemetry to migrate logs from multiple sources. Besides collecting OpenTelemetry data, users can run the Dynatrace OTel Collector to collect all logs from receivers such as syslog, fluentd, or Kubernetes, enrich them, and export them to Dynatrace (or other backends).
  • Mask data for privacy and compliance. Using the Collector to mask data to adhere to privacy and General Data Protection Regulation (GDPR) regulations is another common use case. Users can run the Dynatrace OTel Collector to mask sensitive data next to applications and ensure that no sensitive data leaves the internal network. Users can also filter telemetry data for all signals (traces, metrics, and logs).
  • Ingest and multiplex data. The fully tested Dynatrace OTel Collector distribution also handles ingesting and multiplexing log data for use by different analytics tools and backends, such as Dynatrace for observability, security, and business analytics.

If teams are using the Dynatrace platform as their analytics backend, they benefit from unified analysis of all their observability, security, and business data in full context.

Get started with the Dynatrace OTel Collector

To learn more about the Dynatrace OTel Collector distribution, see the blog Enhance data collection with Dynatrace OTel Collector distribution.

To download the distribution and get started, follow the Collector deployment instructions in the Dynatrace documentation.

The post Dynatrace OTel Collector distribution amplifies OpenTelemetry integration for scalable, production-ready observability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-opentelemetry-collector-for-production-ready-observability/feed/ 0
OpenShift vs. Kubernetes: Understanding the differences https://www.dynatrace.com/news/blog/openshift-vs-kubernetes/ https://www.dynatrace.com/news/blog/openshift-vs-kubernetes/#respond Wed, 07 Jun 2023 19:30:18 +0000 https://www.dynatrace.com/news/?p=58129 OpenShift vs. Kubernetes

Many organizations consider OpenShift vs. Kubernetes for managing containerized apps at scale. What are the differences between OpenShift and Kubernetes?

The post OpenShift vs. Kubernetes: Understanding the differences appeared first on Dynatrace news.

]]>
OpenShift vs. Kubernetes

If you’re evaluating container orchestration software to manage containerized applications at scale, you may be wondering about the differences between OpenShift and Kubernetes. But as you contemplate OpenShift vs. Kubernetes, it’s important to understand what these container orchestration solutions are, how they relate, and their benefits and drawbacks.

A guide to container orchestration software

Container orchestration software automates the administration of containerized workloads and services, greatly reducing the time IT staff spend keeping an application environment running smoothly. Container orchestration allows an organization to digitally transform at a rapid clip without getting bogged down by slow, siloed development, difficult scaling, and high costs associated with optimizing application infrastructure.

As with Kubernetes vs. Docker, OpenShift vs. Kubernetes is a common debate, as they are two of the most widely used container orchestration tools. Although they share many features in common, there are also critical differences between OpenShift and Kubernetes. To put those differences into perspective, let’s look at what Kubernetes and OpenShift are and how they work.

What is Kubernetes?

Kubernetes is an open source container orchestration platform that enables organizations to automatically scale, manage, and deploy containerized applications in distributed environments. According to the Kubernetes in the Wild 2023 report, “Kubernetes is emerging as the operating system of the cloud.” In recent years, cloud service providers such as Amazon Web Services, Microsoft Azure, IBM, and Google began offering Kubernetes as part of their managed services. As a result of these services, organizations can further streamline the administrative overhead associated with application development.

Kubernetes containers are portable across environments, which enables developers to run them on nearly any type of infrastructure, whether in the cloud or locally. This flexibility helps organizations avoid vendor lock-in. Kubernetes also gives developers freedom of choice when selecting operating systems, container runtimes, storage engines, and other key elements for their Kubernetes environments. They can integrate their own applications in the Kubernetes API or use Kubernetes’ own tooling to roll out new features. One major Kubernetes advantage is its self-healing, continually making repairs and addressing failures that affect applications’ integrity.

Kubernetes architecture

That said, Kubernetes has some drawbacks. Its inherent complexity makes observability difficult — especially when used across highly distributed systems. IT teams can’t see into the internal state of Kubernetes containers, so they often collect a wide variety of telemetry data — such as logs, metrics, and distributed traces — to compensate for this lack of visibility.

While these data sources are helpful, they often can’t help IT understand the relationships and context necessary to quickly identify the root causes of application performance issues. As a result, organizations can have trouble transforming at scale, improving critical service-level agreements, and optimizing the user experience.

What is OpenShift?

Like Kubernetes, OpenShift is an open source Kubernetes-based container platform. OpenShift is developed by Red Hat and can run in a variety of environments — both cloud and on premises. In fact, it is a frequent choice for running Kubernetes on premises.

Because it’s based on Kubernetes, OpenShift provides containerization and orchestration of containerized workloads. Like Kubernetes, it allocates resources efficiently and ensures high availability and fault tolerance.

But OpenShift builds from there to provide integrated development tools, CI/CD (continuous integration/continuous deployment) pipelines, and built-in support for popular programming languages, frameworks, and databases. These tools enable OpenShift to support the entire application lifecycle, from development to production, including scaling, rolling updates, and version control. Without having to worry about underlying infrastructure concerns, such as storage, security, and lifecycle management, developers can focus on writing code.

Likewise, Red Hat OpenShift helps organizations administer Kubernetes more efficiently. For example, OpenShift simplifies Kubernetes management tools, giving developers everything they need to manage Kubernetes nodes, as well as the underlying control plane.

In addition, OpenShift provides numerous cloud services and self-managed deployment models to suit various applications and architectures, including the following:

  • OpenShift Container Platform (OCP). This self-managed offering can run on premises or in the cloud.
  • OpenShift Dedicated (OSD). The managed service runs on public clouds such as Amazon Web Services and Google Cloud.
  • Red Hat OpenShift Online (OSO). This fully managed service runs on Red Hat’s public cloud.

Despite its advantages, however, OpenShift also has its limitations. While Kubernetes supports all cloud and Linux distributions, making it widely accessible to organizations using various platforms, OpenShift supports only Red Hat distributions, such as Red Hat Enterprise Linux (RHEL), CentOS, and Fedora. As a commercial solution, OpenShift is also comparatively less flexible than open source Kubernetes, making it less customizable to an organization’s unique requirements.

OpenShift vs. Kubernetes: Weighing the key differences

While Kubernetes and OpenShift are both popular container orchestration platforms, they are used in slightly different ways and offer different features. Some of the key differences include the following:

  • Origin. Originally created by Google, Kubernetes is an open source project managed by the Cloud Native Computing Foundation (CNCF). OpenShift, on the other hand, is an open source Red Hat offering that is built on top of Kubernetes primarily on RHEL operating systems.
  • Ease of use. While Kubernetes offers increased flexibility and powerful features, it can be complex to set up and manage. In contrast, OpenShift provides a simplified, user-friendly interface, with built-in support for CI/CD pipelines.
  • Security. OpenShift has several built-in security features, while Kubernetes relies on the underlying infrastructure and additional tools for security. Additionally, OpenShift runs containers as a non-root user by default and provides additional security policies out of the box.
  • Networking. Kubernetes provides a basic networking model. However, it needs additional tools or plugins for more advanced networking features. OpenShift, on the other hand, includes a more advanced software-defined networking (SDN) solution, which supports network policies for finer control over container communication.
  • Updates and support. Kubernetes has frequent updates, which can sometimes lead to issues such as breaking changes. Red Hat OpenShift offers long-term support versions and commercial support.
  • Integration and extensions. Kubernetes is more of a bare-bones platform. Therefore, it relies on external tools and services for most integrations and extensions. Conversely, as a Red Hat offering, OpenShift provides built-in integration with other Red Hat products and offers a marketplace for third-party extensions.
  • Pricing. Unlike Kubernetes, which is a completely free and open source service, OpenShift has a pricing model for its enterprise version that includes additional features, support, and services.

A lesson in terminology: Kubernetes namespace vs. OpenShift project

It’s OpenShift vs. Kubernetes when it comes to terminology, too. In addition to the aforementioned feature differences, Kubernetes and OpenShift use different terminology, which can be confusing for organizations and practitioners alike.

For example, a namespace in Kubernetes is typically referred to as a project in OpenShift. Despite their similarities, there are a few differences between namespaces and projects, including the following:

  • Access control. In Kubernetes, users manage access control independently from namespaces. In contrast, OpenShift’s projects have a predefined set of permissions for project-level operations. This makes it easier for organizations to control who has access to what within a project.
  • Isolation. While teams use both namespaces and projects to isolate resources within a cluster, OpenShift’s projects provide additional features. These include the ability to limit the amount of resources that all containers can consume within a project.
  • User-friendly. Unlike Kubernetes namespaces, OpenShift projects are more user-friendly. When a user creates a project, for instance, they automatically become the project admin. With Kubernetes namespaces, this does not happen automatically.

Which container orchestration software is right for you?

If your organization needs a container orchestration solution with enterprise-level support and security, OpenShift is the clear choice. OpenShift offers a secure-by-default option to increase security, and its security policies are much stricter than Kubernetes. OpenShift is also a good choice if CI/CD is a priority for your organization.

Additionally, OpenShift is designed to meet the needs of industries with strong compliance and regulatory requirements, such as healthcare or finance. It addresses regulations such as the European Union’s General Data Protection Regulation and the U.S. Health Insurance Portability and Accountability Act.

On the other hand, Kubernetes is a strong option if you need more customization and flexibility, and you have in-house Kubernetes experts who can troubleshoot problems as they arise.

Kubernetes also works on the widest possible range of operating systems and platforms. If you’re a social media or gaming company that places an especially high priority on releasing updates at a rapid pace, Kubernetes may be a better fit.

Automatic and intelligent observability for OpenShift and Kubernetes

Whether you choose OpenShift vs. Kubernetes or vice versa, Dynatrace can make the most of your container orchestration solution. Dynatrace uses AIOps and cloud observability to combine metrics, logs, and traces with topology information, real user experience data, and meta information. With advanced observability of every Kubernetes cluster, pod, and node — and all connections and dependencies they touch — you can quickly pinpoint and solve performance problems as they arise.

Additionally, Dynatrace offers powerful monitoring capabilities for OpenShift, helping you manage costs, automate your operations, and release better software faster.

Whether using OpenShift or Kubernetes, the Dynatrace observability and security platform is the only Kubernetes monitoring system with continuous automation that identifies and prioritizes alerts from applications and infrastructure without changing code, container images, or deployments.

To learn more about how Dynatrace can help you achieve your container orchestration goals, check out our performance clinic, “Kubernetes platform observability with Dynatrace.”

The post OpenShift vs. Kubernetes: Understanding the differences appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/openshift-vs-kubernetes/feed/ 0
Kubernetes OOMKilled troubleshooting: Diagnosing out-of-memory issues automatically https://www.dynatrace.com/news/blog/kubernetes-oomkilled-out-of-memory-troubleshooting/ https://www.dynatrace.com/news/blog/kubernetes-oomkilled-out-of-memory-troubleshooting/#respond Mon, 05 Dec 2022 08:03:12 +0000 https://www.dynatrace.com/news/?p=54991 person shines a light on the Kubernetes logo to discover the root cause of OOMKilled out of memory errors, Kubernetes adoption and Kubernetes survey

Anyone who has struggled with Kubernetes OOMKilled (out of memory) issues knows how frustrating they can be to debug. Kubernetes has a lot of built-in capabilities to ensure your workloads get enough CPU and memory to stay healthy. However, misconfiguration is a common reason why Kubernetes might kill pods despite the workload just doing fine […]

The post Kubernetes OOMKilled troubleshooting: Diagnosing out-of-memory issues automatically appeared first on Dynatrace news.

]]>
person shines a light on the Kubernetes logo to discover the root cause of OOMKilled out of memory errors, Kubernetes adoption and Kubernetes survey

Anyone who has struggled with Kubernetes OOMKilled (out of memory) issues knows how frustrating they can be to debug.

Kubernetes has a lot of built-in capabilities to ensure your workloads get enough CPU and memory to stay healthy. However, misconfiguration is a common reason why Kubernetes might kill pods despite the workload just doing fine and all your Kubernetes nodes still having enough free resources.

Assigning memory resources to pods and containers is a science and an art. Despite good Kubernetes knowledge and best intentions, bad things can, and probably will happen, especially in dynamic software development. While you may have allocated adequate request and limit memory resources for one software version, those settings may not work for another. And when things do go wrong with memory allocation, it’s good to have troubleshooting tips at hand.

This troubleshooting story was brought to me by my friend, Robert Wunderer, founder of CapriSys GmbH.

An out-of-memory issue lurks

Robert runs a multi-tenant e-commerce system on a managed Kubernetes environment. Each tenant gets its own e-commerce site deployed on a shared Kubernetes cluster, isolated through separate namespaces and additional traffic isolation.

The following screenshot gives an overview of some of those namespaces with details of workloads, services, and allocated resources:

screenshot showing separate tenants in OOMKilled out of memory investigation
Each tenant runs in a separate Kubernetes namespace that contains all relevant services: UI, API, backend, reporting, storage, and so on.

A few weeks ago, Robert’s team rolled out a new software version, but things took an unexpected turn. Here’s what happened, the steps Robert took to troubleshoot the problem, and how Dynatrace helps automate the forensics.

OOMKilled discovery: High error rates on backend service

It started on a Friday in the late afternoon, when Robert’s team rolled out a new version across all tenants. There was not much traffic during the weekend, but as Monday came along, Dynatrace started sending alerts about a high HTTP failure rate across almost every tenant on the backend service. The following screenshot shows the list of problems Dynatrace detected.

Screenshot shows high HTTP failure rate in OOMKilled out of memory investigation
Dynatrace automatically alerted on high HTTP failure rates across multiple e-commerce tenants.

Robert first looked through the application logs of the backend service where he noticed intermittent backend failures when contacting the report service. As the name implies, the report service creates various user-defined PDF reports that can be triggered by the user.

Just looking at the application logs for the report service, however, didn’t provide much clarity except that the report-service pod kept restarting with no obvious application logic errors. No exceptions, no error logs, nothing. Both backend and report service are implemented in Java, but Robert found no Java runtime-related issues, either.

Analyzing: Looking for unusual Kubernetes event activity

So Robert analyzed the data and events Dynatrace captured for the Kubernetes cluster responsible for keeping the pods running. He immediately spotted an unusual spike in the number of out-of-memory events (OOMKilled) shown in the following Dynatrace Kubernetes screens.

Screenshot that shows the spike in OOMKilled errors on the report-service
Dynatrace ingests Kubernetes events and assigns them to the monitored entities. Easy to spot the spike in OOMKilled errors on the report service.

This data confirmed that the system was killing off the report service pods for lack of memory. During the container restart, the backend service could not connect to the report service, which led to the errors.

The memory needed for running those reports varies depending on the number and size of the requested report. The question was whether the problem was due to actual memory pressure or whether it was caused by too-tight resource settings. Did somebody change the memory settings or did the memory requirements of the new version change?

Dynatrace automatically detects any deployment specification change (such as replicaset, version, limits, and so on) shown in the list of Kubernetes events. In the following screenshot, you can see that Dynatrace automatically detected the deployment specification version change, which went from 3.0 to 4.0. But that’s about it. No other change was detected.

Screenshot that shows the spec change events before the OOMKilled errors spiked
Deployment specification change events at the time before the OOMKilled errors spiked showed the version update from 3.0 to 4.0.

Besides the version update, nothing has changed. So why were the pods killed with an OOMKilled error? Was there a general memory shortage on the Kubernetes cluster nodes? Or did version 4.0 of the report service just need more memory than the previous 3.0 version?

Out-of-memory root cause: Wrong settings or shortage of resources?

Because the problem was happening with almost all the report service pods and Dynatrace reported no problems with memory on any of the Kubernetes nodes, Robert already suspected that the memory behavior of the new version has changed.

Looking at the resource analyses confirmed the suspicion. The following chart shows how memory usage rises to almost the defined limit of 400MB before the pod is killed off and restarted. Historical data shows this didn’t happen with the old version.

Screenshot that shows pod memory metrics indicating how k8s kills the pod when it reaches out of memory limit
Analyzing the pod memory metrics shows how Kubnernetes kills the pod when it reaches the memory limit.

The fix: Adjust the memory settings to avoid OOMKilled errors

Armed with that knowledge, the solution was simple: Robert increased the pod’s memory limit from 400MB to 600MB so that reports didn’t run into the out-of-memory shortage.

Screenshot that shows proper sizing of k8s resources to avoid out of memory error
Lesson learned: It’s important to do proper sizing of your Kubernetes resources on a deployment level.

Lesson learned: Avoid out of memory issues by regularly testing memory allocation

In my conversation with Robert, he told me they mainly do functional testing before promoting a new version to production. Now they’re investing in more load, scalability, and memory testing with the goal of better understanding capacity requirements.

If they had run more tests with varying report sizes, they would have identified the changed memory behavior. This could have resulted in either fixing a memory problem or simply identifying the proper limits for smooth production operations.

An automatic safety net for discovering out-of-memory and other errors

With Dynatrace in place, Robert could immediately pinpoint the root cause of the OOMKilled errors, which confirmed his suspicions. This instant answer spared his team a lengthy troubleshooting process so they could get right to implementing more comprehensive testing.

I want to thank Robert for sharing this story. Because Kubernetes is becoming the default platform for most of our workloads, it’s important to know what can go wrong. It’s also important to educate others about how Kubernetes handles resource settings and how we can use observability data to detect the root cause and take corrective actions.

If you want to learn more about Kubernetes, I also highly recommend the following resources:

To try Dynatrace yourself, sign up for your own Dynatrace SaaS Trial.

Kubernetes in the wild report 2025

Uncover global Kubernetes adoption trends, cost-optimization strategies, and key tools driving innovation for thousands of organizations worldwide. This report highlights global trends in the technology’s adoption and usage in production environments from thousands of organizations across diverse industries.

Kubernetes clusters hosted in the Cloud

The post Kubernetes OOMKilled troubleshooting: Diagnosing out-of-memory issues automatically appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/kubernetes-oomkilled-out-of-memory-troubleshooting/feed/ 0
How to collect Prometheus metrics in Dynatrace https://www.dynatrace.com/news/blog/how-to-collect-prometheus-metrics-in-dynatrace/ https://www.dynatrace.com/news/blog/how-to-collect-prometheus-metrics-in-dynatrace/#respond Tue, 16 Nov 2021 10:00:12 +0000 https://www.dynatrace.com/news/?p=47136 What is Prometheus and 4 challenges for enterprise adoption

Dynatrace has recently extended its Kubernetes operator by adding a new feature, the Prometheus OpenMetrics Ingest, which enables you to import Prometheus metrics in Dynatrace and build SLO and anomaly detection dashboards with Prometheus data. Here we’ll explore how to collect Prometheus metrics and what you can achieve with them. What is Prometheus? Before we […]

The post How to collect Prometheus metrics in Dynatrace appeared first on Dynatrace news.

]]>
What is Prometheus and 4 challenges for enterprise adoption

Dynatrace has recently extended its Kubernetes operator by adding a new feature, the Prometheus OpenMetrics Ingest, which enables you to import Prometheus metrics in Dynatrace and build SLO and anomaly detection dashboards with Prometheus data. Here we’ll explore how to collect Prometheus metrics and what you can achieve with them.

Dynatrace dashboard - Prometheus
The dashboard was built during my recent Performance Clinic on k8s monitoring at Scale with Prometheus and Dynatrace.

What is Prometheus?

Before we get into how to collect Prometheus metrics in Dynatrace, let’s first outline what Prometheus is and what it does.

Prometheus is an open-source software toolkit used for event monitoring and alerting. It records real-time metrics in a time series database with flexible queries and real-time alerting.

How to collect Prometheus metrics

Since its launch in 2012, Prometheus has become the standard technology to collect metrics in a Kubernetes cluster. Prometheus pulls (scrapes) real-time metrics from application services and hosts by sending HTTP requests on Prometheus metrics exporters. It then compresses and stores them in a time-series database on a regular cadence.

The new Dynatrace Kubernetes operator can collect metrics exposed by your exporters. This means Dynatrace isn’t collecting the metrics on the Prometheus server, but directly at the source of truth – the exporters.

Additionally, you don’t have to worry about scaling the Prometheus infrastructure because it doesn’t even have to be collected by the Prometheus server.

Once the data is ingested by Dynatrace, users can then take advantage of all Dynatrace features:

  • Manage zones to control access to your data
  • SLO and anomaly detection rules to handle your alerts directly in Dynatrace
  • Dashboards to visualize metrics in context including Dynatrace metrics, Prometheus, custom metrics, logs, and so on.

How do you configure the Prometheus OpenMetrics Ingest?

For Dynatrace to take advantage of all the metrics exposed by the various Prometheus exporters, simply add a few annotations to your exporter’s deployment file. The Dynatrace Kubernetes operator automatically scrapes metrics on pods and services that have Dynatrace annotations. Dynatrace needs to know:

  • The HTTP endpoint of your exporter that exposes Prometheus metrics
  • The port
  • Any required authentication settings

With that in mind, here’s a step-by-step guide on how to:

  • Identify the scraping configuration of your exporter
  • Add the Dynatrace annotation
  • Know where to find your Prometheus metrics ingested by Dynatrace.

1. Identify the scraping configuration of your exporter

Every Prometheus exporter can have different settings to collect the metrics exposed. These are Port, Path, and Protocol (HTTP or SSL). Those settings are crucial to allow Prometheus (or Dynatrace in our case) to scrape the metrics.

Port

How do you get the port of your exporter? You can find the information related to the port in the description of your pod. If you run the command shown below, you’ll get all the information related to your pod:

Command to display the details related to our pod
Command to display the details related to our pod

Example of output of the command:

Example of output of the command
In this example, we can clearly see that the exporter listens on port 9100.

Path

Most of the exporters are exposing the metrics on the path: /metrics. But /metrics is not standard, so it means exporters can use any path to expose their data.

We highly recommend reading the documentation of your exporter to get the right path.

Here is a website listing the various exporters:  PromCat.io – A resource catalog for enterprise-class Prometheus monitoring

Protocol:

By default, Prometheus utilizes the HTTP protocol to scrape the metrics out of the exporter.

But to add extra security, you can deploy NGNIX to enable the SSL encryption and also add an extra authentication policy.

2. How to test the configuration?

To ensure the correct configuration of the Dynatrace annotations, we recommend testing the settings by sending a request to the exporter’s HTTP endpoint.

Once you have identified the port of your exporter you will simply have to create a port forward:

Comand line port forward

Then direct your browser to http://localhost:9090/metrics (see my example) and see the metrics in the Prometheus format.

You can also test the settings by deploying a pod with Curl installed to test your request directly within your cluster.

3. How to add the Dynatrace annotation?

Similarly to Prometheus, Dynatrace needs to be able to “discover” the pods exposing Prometheus metrics. To accomplish this, add an annotation on your pods/services deployment file.

Comand line pods/services deployment file

If you are not comfortable with changing the deployments of your exporters, you can also create a new “fake” service with the required annotations and settings to select the pods to scrape.

code example create 'fake'service

Keep in mind that Dynatrace will not be able to ingest Prometheus metrics if the annotations are not configured properly.

Where can I see my Prometheus metrics in Dynatrace?

You can see the Prometheus metrics in:

  • Metrics
  • Or the Data explorer.
Data explorer with Prometheus metrics Dynatrace
Example of a query that will calculate the ratio of running pods (with the metrics collected from the kubestate metrics)

The Data Explorer will allow you to build queries on your Prometheus data.

One of the advantages of the Prometheus ingest is that all the Prometheus labels will be available as dimensions in Dynatrace.

Dimensions allow you to achieve various operations such as:

  • Grouping and filtering in the Data explorer
  • Defining management zones
  • Create anomaly detection rules

Once the metrics are ingested, we will of course be able to analyze our Prometheus metrics in Dynatrace, but the other great advantage of Dynatrace is the ability to create anomaly detection rules based on your “queries”.

Prometheus metrics Dynatrace sceenshot

How can we automate the deployment and the configuration of Dynatrace?

As a Prometheus user, you will probably expect to be able to deploy new exporters, build graphs and define alerts automatically through a CI process or a script.

Dynatrace has a tool named Monaco (Monitoring as Code) that allows you to automate the configuration of your Dynatrace tenant.

Monaco is a command-line tool that interacts with the configuration API. This tool enables you to implement

  • Dashboarding as code
  • Creation of management zones
  • Alerting as code
  • And much more

With Prometheus OpenMetrics ingest and Monaco, you can implement the following pipeline:

Prometheus OpenMetrics ingest and Monaco pipeline Dynatrace

This automates:

  • The deployment of your new exporter and the “fake” service that will add the Dynatrace scraping annotation
  • The configuration of management zones
  • The creation of the new dashboard and anomaly detection rule.

Getting started with Prometheus

Over the years Prometheus has become an industry standard. Most software technologies currently on the market are exposing observability metrics related to their product (CI/CD, network appliances, databases).

As a cloud-native engineer, the Prometheus OpenMetrics ingest capability is amazing news because it allows us to fully take advantage of these external metrics within Dynatrace to observe our IT environment, our CI/CD process, our backup process, and more.

Once Dynatrace ingests the external metrics, you will be able to take advantage of all the great features and proactively avoid customer-facing breakdowns in case of an issue.

Similar to the configuration of Prometheus, adding external metrics means you can add the right annotation on your Pods/Services deployment files.

By importing Prometheus metrics into Dynatrace you also gain :

  • User permissions management of top of your Prometheus metrics
  • Significant reduction of maintenance tasks related to Prometheus infrastructure
  • Full advantage of all Dynatrace features (notify your teams in case of problems, define dashboards, SLOs, etc.) on the top of Prometheus native metrics.

Prometheus is a powerful tool, but it also comes with a few challenges. To learn more about these challenges, check out my other blog – What is Prometheus and Four Challenges for enterprise adoption.

The post How to collect Prometheus metrics in Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/how-to-collect-prometheus-metrics-in-dynatrace/feed/ 0
Kubernetes workload troubleshooting with metrics, logs, and traces https://www.dynatrace.com/news/blog/kubernetes-workload-troubleshooting-with-metrics-logs-and-traces/ https://www.dynatrace.com/news/blog/kubernetes-workload-troubleshooting-with-metrics-logs-and-traces/#respond Thu, 21 Oct 2021 14:48:52 +0000 https://www.dynatrace.com/news/?p=46712 Dynatrace employees

There’s no lack of metrics, logs, traces, or events when monitoring your Kubernetes (K8s) workloads. But there is a lack of time for DevOps, SRE, and developers to analyze all this data to identify whether there’s a user impacting problem and if so – what the root cause is to fix it fast. At Dynatrace […]

The post Kubernetes workload troubleshooting with metrics, logs, and traces appeared first on Dynatrace news.

]]>
Dynatrace employees

There’s no lack of metrics, logs, traces, or events when monitoring your Kubernetes (K8s) workloads. But there is a lack of time for DevOps, SRE, and developers to analyze all this data to identify whether there’s a user impacting problem and if so – what the root cause is to fix it fast.

At Dynatrace we’re lucky to have Dynatrace monitor our workloads running on K8s. One of those workloads is Keptn, a CNCF project Dynatrace is contributing to, that we use internally for different SLO-driven automation use cases.

Dynatrace Davis, our deterministic AI, recently notified our teams about a problem in one of our Keptn instances we just recently spun up to demo our automated performance analysis capabilities orchestrated by Keptn. I was pulled into that troubleshooting call and started taking notes and screenshots so I can share how easy it is to troubleshoot the Kubernetes workload with our engineers and you – our readers – on this blog post.

It started with the Problem card Davis opened because of a 33% increase in failure rate in the workload called mongodb-datastore.

Dynatrace automatically baselines all service endpoints of all deployed workloads in k8s and alerts on abnormal behavior such as a jump in failure rate
Dynatrace automatically baselines all service endpoints of all deployed workloads in k8s and alerts on abnormal behavior such as a jump in failure rate

This mongodb-datastore provides several internal API endpoints to fetch/update data in the actual MongoDB instance. What’s great for our engineers who are responsible to operate Keptn is that alerting happens automatically, without relying on us to define custom thresholds. This is all thanks to Dynatrace’s automatic adaptive baselining.

How the baselining identified the problem can be easily seen with a single click from the problem to the service response time overview as shown next:

Response time, failure rate or throughput are automatically baselined for every service and service endpoint. In this case – Failure Rate jumped to an unusual high value compared to the automatic baseline
Response time, failure rate or throughput are automatically baselined for every service and service endpoint. In this case – Failure Rate jumped to an unusual high value compared to the automatic baseline

What’s even better than Davis detecting the increase in failure rate is that Davis automatically points to the root cause, which Dynatrace picked up from automatically captured container logs. A single click brought us to the Log screen – automatically filtered to the logs captured in that mongodb-datastore during that timeframe:

Dynatrace automatically captures all container logs and shows them in context of a detected problem. Like this unhandled exception leading to a crash of the process
Dynatrace automatically captures all container logs and shows them in context of a detected problem. Like this unhandled exception leading to a crash of the process

If you take a closer look at the screenshot above it’s easy to spot the root cause; it was an unhandled error condition in the code that was waiting and processing feedback from the MongoDB instance. The problem was also reported back to the Keptn team via GitHub issue mongodb-datastore: Panics, meaning the team could not only detect the issue fast but also had everything they needed to react fast and immediately fix the problem.

Dynatrace has even more details for the development teams

Just from spending two minutes looking at the data Davis put in front of us, we knew the Impact and Root Cause of the high error rate (including the line of code). I went ahead and took additional screenshots that I sent over to the engineers on top of the direct links to the Dynatrace screens so that they can do their own analysis. Screenshots are however always great as I think they are interesting for the engineers and make my life in writing those blogs easier 😊

One of those screenshots is from one of the PurePaths (=Distributed Traces) that captured the problem. Not only do we have the detailed log, but we also know the API endpoint was the HTTP GET /event.

PurePaths (=Distributed Traces) give additional context for developers such as timing, endpoints, caller information …
PurePaths (=Distributed Traces) give additional context for developers such as timing, endpoints, caller information …

There’s so much more Dynatrace provides than what’s shown here, but I wanted to keep this blog short and sweet and focus on that one story.

If you’re interested in learning more, I recommend you check out these articles:

As always – these blog posts wouldn’t be possible if my colleagues wouldn’t share those stories. Thanks a lot Sergio Hinojosa and Maria Rolbiecka for your hard work on keeping this Keptn instance up and running! And a big Thank You to Florian Bacher and Bernd Warmuth from the Keptn development team who went ahead and fixed that issue within a day – now that’s a fast turnaround 😊

Learn more about Kubernetes observability for SREs with this YouTube tutorial by Henrik Rexed.

The post Kubernetes workload troubleshooting with metrics, logs, and traces appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/kubernetes-workload-troubleshooting-with-metrics-logs-and-traces/feed/ 0