Infrastructure Archives | Dynatrace news https://www.dynatrace.com/news/category/infrastructure/ The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Wed, 01 Jul 2026 09:48:05 +0000 en hourly 1 How Dynatrace supercharged log observability in 2025 https://www.dynatrace.com/news/blog/how-dynatrace-supercharged-log-observability-in-2025/ https://www.dynatrace.com/news/blog/how-dynatrace-supercharged-log-observability-in-2025/#respond Thu, 15 Jan 2026 17:18:49 +0000 https://www.dynatrace.com/news/?p=72456 Dynatrace Logs icon

Large enterprises such as Western Union, Vodafone, and United Airlines are ditching legacy log solutions in favor of a single, unified observability platform that delivers real-time insights and scalability at the petabyte level, as you’ll hear firsthand from them at Perform 2026. In this blog post, we’ll look back at the log-focused Dynatrace product releases […]

The post How Dynatrace supercharged log observability in 2025 appeared first on Dynatrace news.

]]>
Dynatrace Logs icon

Large enterprises such as Western Union, Vodafone, and United Airlines are ditching legacy log solutions in favor of a single, unified observability platform that delivers real-time insights and scalability at the petabyte level, as you’ll hear firsthand from them at Perform 2026.

In this blog post, we’ll look back at the log-focused Dynatrace product releases of 2025, while keeping in mind the three benefits that customers love most about Dynatrace:

  1. Fast log onboarding with unified ingestion from any source
  2. It’s easy to get started, yet powerful for your daily work
  3. Productivity boosts with Davis AI

Boost productivity with Davis AI

The promise of “Logs in Context” is simple:
Find the right log line at the right time, automatically and powered by AI.

The magic of Dynatrace is not a single feature or hyped AI. It’s the sum of many Dynatrace capabilities that comprise the foundation of the Dynatrace platform: Grail®, Smartscape®, Davis® AI, OpenPipeline®, and many others, that come at no extra cost, providing the automation and assistance you need.

Easily identify root causes and create tickets using the Dynatrace Problems app and logs.
Figure 1. Easily identify root causes and create tickets using the Dynatrace Problems app and logs.

If you aren’t yet using Dynatrace for your logs, stop stitching together clues across tools and say goodbye to manual swivel chair ops:

  • Logs in context: The right log lines appear automatically within the workflow or Dynatrace app you’re using. Whether that’s troubleshooting a service, reviewing Kubernetes node health, or investigating performance incidents of Infrastructure or cloud native apps.
  • Free of charge: Every in-context query, including surrounding logs, is now zero-rated (non-billable) when you view logs inside these Dynatrace core apps: Clouds, Infrastructure & Operations, Services, and Distributed Traces. While these apps don’t generate query consumption, ingestion and retention consumption are billed individually. We’re delivering the logs you need to take action – instantly, efficiently, and automatically correlated.
  • Leverage the power of Dynatrace Davis AI: With Dynatrace, features like “Explain logs” dramatically shorten time to action. Our customers report that their teams can more easily understand the possible causes and impacts of incidents without having to manually search for error codes in logs on Google.
  • By leveraging Davis AI, Workflow Automation, and integrations such as our ServiceNow partnership, customers can dramatically reduce the number of incidents; one of our customers reported reducing MTTI by 90%.

AI summaries are available across the Dynatrace platform and MCP server.

Explore logs, expand log messages, and comprehend them faster using the “explain log” AI feature.
Figure 2. Explore logs, expand log messages, and comprehend them faster using the “explain log” AI feature.

With Dynatrace, observability is not limited to cloud native apps. These features work seamlessly across cloud native, on-premises, hybrid, and traditional IT stacks. So, whether you’re on Kubernetes, a Mainframe, or an AWS Lambda function, the experience is the same.

Effortless for everyone, powerful for experts

Once your logs are ingested, you need to be able to understand them. This is where our Logs app shines for both new and expert Dynatrace users.

Pre-defined and admin-curated views boost productivity

Earlier this year, we improved the simplicity of applying complex and advanced queries with new data segmentation and advanced filters.

Using segments, admins and power-users can provide reusable and pre-scoped filters. When paired with dynamic variables, users can easily modify filter conditions.

Simultaneously, we continued enhancing the Logs app to provide advanced click-to-filter capabilities in various areas, like pinning frequent queries and filters:

  • Filter field: Suggest attributes, operators, and entities
  • Facets: Gain a quick understanding of patterns and groups, or build queries
  • Advanced filtering: Intuitive click-to-filter side pane, including JSON-structure log support with nesting
Combine segments and facets to create a pre-filtered view
Figure 3. Combine segments and facets to create a pre-filtered view

JSON‑structured log handling

Log messages aren’t always clean. A field might be hidden inside a nested message attribute or buried three levels deep in nested JSON.

Dynatrace log handling:

  • Detects and normalizes JSON.
  • Exposes nested fields in the UI without manual mapping.
  • Provides human-readable log messages in the results across all apps that use logs.

This way, you and your users can focus on analysis, not plumbing and normalizing logs.

Free text search surfaces the content you're looking for instantly, with human-readable results, even for JSON-structured log records
Figure 4. Free text search surfaces the content you’re looking for instantly, with human-readable results, even for JSON-structured log records

Correlation at scale

With Traces on Grail, your traces are automatically correlated in context with surfaced logs within the Distributed Tracing app, including associated exceptions.

The value you and your teams gain

If you’re accustomed to working with traces, you can continue using your troubleshooting routine and easily navigate from traces to logs and error exception messages. If you prefer to start your work by focusing on logs, you can achieve the same outcome.

The Dynatrace Distributed Tracing app automatically links logs with traces or spans.
Figure 5. The Dynatrace Distributed Tracing app automatically links logs to traces or spans.

Remember, Logs in context are free with the Distributed Tracing app!

Fast log onboarding with unified ingestion from any source

You want all your logs, and you want them fast. You don’t want to wrestle with YAML files, forward scripts, or configure custom collectors.

Centralized configuration, self-service management, and enabling teams with granular permissions to collect and ingest logs—these are what customers asked for:

  • OneAgent + Journald – Enhanced capabilities for automatically capturing logs on Linux machines with a single, centralized, configured agent: Dynatrace OneAgent®. Just deploy and watch the logs magically appear in your tenant.

Kubernetes logging made easy – The Dynatrace Kubernetes Logs Module gives you complete visibility without requiring OneAgent to operate in Full-Stack mode or to configure OTel manually.

Onboarding your Kubernetes cluster and logs using the Log Onboarding Wizard.
Figure 6. Onboarding your Kubernetes cluster and logs using the Log Onboarding Wizard.
  • Log Onboarding Wizard – To further simplify the onboarding experience, we’ve introduced a new wizard across several apps. When logs are missing, or you manually launch the wizard, it provides guided steps to onboard your logs, including creating an API key.
If you already have a standardized intake process in place for your teams, simply don’t provide one or all of the required permissions. Then your users won't be able to see the wizard or onboarding recommendations.
Figure 7. If you already have a standardized intake process in place for your teams, simply don’t provide one or all of the required permissions. Then your users won’t be able to see the wizard or onboarding recommendations.

Scale that never breaks

You can ingest up to 1 PB of logs per day per tenant, which should eliminate most sizing or scaling headaches. This bandwidth is part of the Dynatrace SaaS magic: Dynatrace Grail stores and processes everything in an indexless manner and using schema on-read. At the same time, OpenPipeline® routes the telemetry according to your rules and requirements defined in the pipelines.

  • Thousands of pipelines and self-service: Every pipeline has fine-grained permissions, so teams can create isolated pipelines and self-service onboarding, processing, and routing to buckets for retention.
  • 120+ parsing processors – From JSON normalization to custom field extraction, you can assign a processor to a pipeline and let OneAgent do the matching magic for you. Are you using OpenTelemetry or Cribl? Matching conditions or technology attribution offers you the same experience, regardless of the log source.
    10 MB log records – In our Go big with Dynatrace blog post, we discussed why large log records aren’t an anomaly and how customers benefit from out-of-the-box support for large log records.
    Figure 8. 10 MB log records – In our Go big with Dynatrace blog post, we discussed why large log records aren’t an anomaly and how customers benefit from out-of-the-box support for large log records.

With Dynatrace, your log ingestion stays ahead of your growth curve, no matter how many new sources you add.

Keep your costs predictable

Large enterprises often need to charge back to internal business units. We’ve introduced increased flexibility for existing features related to chargebacks:

  • Retain with Included Queries: Configurable on the individual bucket level, and seamlessly combinable with the established usage-based IRQ model.
  • Cost Allocation: Attribute your logs, metrics, and traces with business‑unit and product labels. This supports your FinOps efforts, as recently discussed in our blog post, Cost Allocation for Logs.

Best Practices: Not everything we delivered in 2025 was a product enhancement. We’ve also delivered a new best practices section in our product documentation, based on field feedback from pre-sales, post-sales, and support teams.

If you prefer to watch a webinar recording instead of reading, we recorded a video that walks you through all the best practices detailed in this blog post.

Dynatrace YouTube Series  | Optimize your logs: Save money and boost performance

Ready to get started?

Let’s make observability effortless, not overwhelming.

Your team can spend less time chasing logs and more time delivering value. Dive in today and experience the power of an observability platform that was built for the future of IT.

If you’re using Dynatrace SaaS with a DPS contract, all the features mentioned in this blog post are available to you. If you’re not, why not start a free trial today and experience the value yourself?

Resources

Dynatrace University – Free training that covers everything from basic log ingestion to advanced analytics.

Dynatrace Playground – Our free sandbox tenant with sample log files, ready to explore log

State of Log Management 2026 – Download the report to explore benchmark data on how AI workloads are exploding log volume and costs, and why unified observability is now essential for reliable, trustworthy AI.

The post How Dynatrace supercharged log observability in 2025 appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/how-dynatrace-supercharged-log-observability-in-2025/feed/ 0
Accelerate SNMP network device observability with Dynatrace Discovery & Coverage https://www.dynatrace.com/news/blog/accelerate-snmp-network-device-observability-with-dynatrace-discovery-coverage/ https://www.dynatrace.com/news/blog/accelerate-snmp-network-device-observability-with-dynatrace-discovery-coverage/#respond Tue, 06 Jan 2026 20:23:04 +0000 https://www.dynatrace.com/news/?p=72365 SNMP Autodiscovery

When onboarding network devices for observability, challenges often arise related to inconsistent or partial monitoring coverage or inefficient processes. Such challenges make it difficult to ensure that all devices are properly and uniformly monitored and can provide actionable insights. Managing network devices at scale exacerbates the problem, as organizations contend with thousands of devices from […]

The post Accelerate SNMP network device observability with Dynatrace Discovery & Coverage appeared first on Dynatrace news.

]]>
SNMP Autodiscovery

When onboarding network devices for observability, challenges often arise related to inconsistent or partial monitoring coverage or inefficient processes.

Such challenges make it difficult to ensure that all devices are properly and uniformly monitored and can provide actionable insights. Managing network devices at scale exacerbates the problem, as organizations contend with thousands of devices from a diverse range of vendors. This can have an adverse impact on your ability to maintain and troubleshoot your networks.

Network monitoring tools often lack integration with the rest of the infrastructure, making it even more challenging to analyze network monitoring data in context.

To achieve consistent, end-to-end monitoring, you need a tool that allows deterministic network observability and automatically contextualizes all ingested data. Let’s take a look at how Dynatrace does this.

Simplified and accelerated network monitoring

When onboarding network devices to an observability platform, it’s not uncommon to spend a disproportionate amount of time configuring network observability. Despite spending valuable time in the process, the uncertainty that some devices might be overlooked remains. Furthermore, in today’s dynamic environments, devices can be added or removed at a moment’s notice and at a rapid pace. This creates a burden for the operations team to constantly prove that the level of observability is at the right level.

To gain a comprehensive overview of the state of observability across all your environments, Dynatrace has expanded the Discovery & Coverage application to include network device monitoring capabilities. This ensures that there are no blind spots in this domain and that no network device is left unmonitored.

The Discovery & Coverage app achieves that by scanning predefined IP ranges or subnets. Based on the outcome of the discovery process, the journey continues by offering the possibility to automatically onboard the right device extension and create Network Availability monitors. Statistics are provided for both extension and availability coverage, providing tangible data to measure the completion of the observability enablement process.

Manage network devices at scale across distributed environments

SNMP (Simple Network Management Protocol) provides a standardized framework for monitoring and managing devices on IP networks. Its simplicity, scalability, and compatibility with a wide range of hardware make it an ideal choice for network management across diverse environments.

However, managing and monitoring SNMP across devices from multiple vendors in large, distributed, or even siloed networks can become cumbersome. Additional complexity is introduced when various teams own and manage their load-balancing devices or application firewalls and core switches individually.

Historically, IP Address Management (IPAM) tools have been effective at mapping entire IP networks, but they struggle to leverage the observability potential of discovered endpoints. Health information on SNMP devices is often isolated, and discovered devices are not placed in the correct context or topology, thereby failing to fulfill the goal of the discovery process: associating SNMP device data with the IT environments where the devices reside. This is where Dynatrace excels.

Leverage the power of the Dynatrace platform for your SNMP devices

Driving tool consolidation and integrating auto discovery and monitoring into your observability solution reduces costs and boosts operational excellence. Additionally, using the Discovery & Coverage app for auto discovery of networking devices means you’ll spend less time manually tracking devices and reduce the chance of lacking the right extension or using the wrong extension for a device.

The Dynatrace end-to-end observability approach makes your discovered network devices available in Dynatrace Grail® and the Dynatrace platform, allowing you to build innovative functionality while simultaneously reducing tool sprawl.

Your discovered devices appear in the Infrastructure and Operations app as network devices, with essential properties automatically populated. Dynatrace and vendor-provided extensions offered in Dynatrace Hub can enhance the observability level and insights of your individual devices. Further, you’re automatically notified when predefined thresholds are exceeded or when anomalies are detected.

Your network devices will function as part of an integrated network, not as standalone entities. For Dynatrace SaaS customers, network devices are readily available in Grail and can be queried using Dynatrace Query Language (DQL) to create value-added insights in the Davis Anomaly Detection app or to create automations and workflows, such as auto-generated tickets for external systems or notifications sent to teams via Slack or PagerDuty.

Get started with network device autodiscovery in Dynatrace

The SNMP autodiscovery capability is provided by the Discovery & Coverage app. Open the app, navigate to the Network coverage tab, and select Configure scanning.

SNMP network device observability configuration in Dynatrace

Edit SNMP network device observability configuration in Dynatrace

Once a configuration is activated, the discovery process begins.

The SNMP Autodiscovery extension regularly scans your IPv4 and IPv6 address lists, ranges, or subnets for devices with SNMP agents. If devices with matching SNMP parameters are present, they will be added to the environment.

Coverage reports for each device configuration provide an overview of the level of observability for each discovered device, answering questions such as: Is each device polled with the correct extension? Or, is the Network Availability Monitor installed on the management interface?

SNMP network device observability in Dynatrace

The Discovery & Coverage app allows you to add a device type-matching extension in bulk. The Poll icon opens the following windows, enabling the specific extensions with one select.

SNMP network device observability poll device in Dynatrace

Additionally, you can centrally configure Network Availability Monitors (NAM) to continuously probe the health and availability of your devices. This ranges from simple ICMP ping tests to advanced probes, depending on your requirements and setup.

SNMP network device observability create a ping monitor in Dynatrace

Take the first step to simplifying device onboarding

SNMP network device autodiscovery is a powerful tool for network administrators. It simplifies network management, improves efficiency, and allows for easy scalability. By leveraging this feature, organizations today ensure their networks are always accurately represented and observed.

Remember, a well-managed network with thorough observability and health monitors is the backbone of any successful organization. So, embrace the power of SNMP autodiscovery and take your network management to the next level!

If you have already deployed network observability using the extension apps, you need to download and open the Discovery & Coverage app. By configuring autodiscovery, you’ll ensure that no part of your network remains unattended and that you’ve configured the right level of observability in your environment. From now on, the Dynatrace network device onboarding process is greatly simplified and accelerated.

If you’re not yet a Dynatrace customer, consider starting a free trial, opening the Discovery & Coverage app, and navigating to the Network coverage tab. You can then immediately discover SNMP devices in one of your environment’s IP networks and add the right level of observability to the devices.

The post Accelerate SNMP network device observability with Dynatrace Discovery & Coverage appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/accelerate-snmp-network-device-observability-with-dynatrace-discovery-coverage/feed/ 0
Get the most from Network Availability Monitoring on Dynatrace Managed https://www.dynatrace.com/news/blog/network-availability-monitoring-on-dynatrace-managed/ https://www.dynatrace.com/news/blog/network-availability-monitoring-on-dynatrace-managed/#respond Tue, 04 Feb 2025 16:00:27 +0000 https://www.dynatrace.com/news/?p=67633 abstract spheres and columns representing hybrid cloud and hybrid kubernetes networks

Network monitoring software can detect slow traffic or component failures due to internal network issues or network connection issues. Network Availability Monitoring (NAM) is a feature of network monitoring, and it ensures the reliability and performance of your IT infrastructure. It allows you to monitor the availability of remote hosts or services over the network when an HTTP/HTTPS endpoint isn't available. Dynatrace Managed extends its synthetic monitoring capabilities to include NAM, providing comprehensive insights into the health of your network and services. This blog post explores the problem of network availability, how Dynatrace addresses it, the specific functionalities of NAM, and what to expect in future updates.

The post Get the most from Network Availability Monitoring on Dynatrace Managed appeared first on Dynatrace news.

]]>
abstract spheres and columns representing hybrid cloud and hybrid kubernetes networks

Expectations for network monitoring

In today’s digital landscape, businesses rely heavily on their IT infrastructure to deliver seamless services to customers. However, network issues can lead to significant downtime, affecting user experience and business operations. Traditional monitoring tools often fall short of providing deep insights into network layers, leaving gaps in understanding the root causes of performance issues. The market demands a robust solution that can monitor applications and the underlying network infrastructure to ensure end-to-end availability and performance.

The Dynatrace approach

Dynatrace addresses these challenges by extending its synthetic monitoring capabilities to include Network Availability Monitoring. NAM allows you to monitor the availability of remote hosts, network devices, and services over the network, even when HTTP/HTTPS endpoints are unavailable. You can utilize the same private synthetic locations you’ve been using to execute HTTP and Browser monitors and run ICMP, TCP, and DNS monitors. By integrating NAM with powerful Dynatrace Davis® AI, you gain 24/7 insights into the health of your network, enabling proactive identification and resolution of issues before they impact your business.

Figure 1. Overview of a Network Availability monitor related to a specific problem thanks to Davis AI
Figure 1. Overview of a Network Availability monitor related to a specific problem thanks to Davis AI

Key features of Network Availability Monitoring in Dynatrace Managed

  • Three types of monitors are available: ICMP (to check the reachability of network devices and hosts), TCP (to verify the availability of specific services running on your network), and DNS (to ensure your DNS services are resolving correctly)
  • All monitors are executed from the same private locations you’re already using to monitor your internal applications with HTTP and browser monitors.
  • NAM integrates seamlessly with the Dynatrace Problems page in the web UI. These synthetic tests are integrated into Dynatrace Managed, allowing you to configure and manage them easily through the Dynatrace web UI or API. The results are combined with other monitoring data to provide a comprehensive view of your IT environment.

In Dynatrace Managed, you can easily create a new Network Availability Monitor by navigating to your list of available network monitors at Digital Experience > Synthetic > List of available Network Monitors and selecting Create a Network Availability Monitor. Configure your Network Availability Monitor according to your use case and save your changes. For more details, please refer to Dynatrace Documentation.

Figure 2. Dynatrace Managed UI for configuring a new Network Availability monitor
Figure 2. Dynatrace Managed UI for configuring a new Network Availability monitor

Next steps

Dynatrace Managed NAM allows the monitoring of remote hosts, network devices, and services through synthetic tests, allowing around-the-clock insights into network health. Future updates may include advanced analytics, extended support for additional network protocols, or integration with other Dynatrace features.

Check out our Dynatrace NAM documentation to learn more about its functionality and how to use it to address your use cases.

For further information, please also consult our newly established Network Availability Monitoring FAQ community post.

The post Get the most from Network Availability Monitoring on Dynatrace Managed appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/network-availability-monitoring-on-dynatrace-managed/feed/ 0
Analyze query performance: The next level of database performance optimization https://www.dynatrace.com/news/blog/analyze-query-performance-the-next-level-of-database-performance-optimization/ https://www.dynatrace.com/news/blog/analyze-query-performance-the-next-level-of-database-performance-optimization/#respond Mon, 04 Nov 2024 21:15:16 +0000 https://www.dynatrace.com/news/?p=66490 Data security graphic

With the recently released version of the Databases app, Dynatrace allows you to monitor your databases from a query perspective: Quickly find all heavy queries that consume your database resources Monitor your query metrics and resource utilization Understand how your queries are executed through the Query Execution Plan Optimize application performance: The importance of database analysis […]

The post Analyze query performance: The next level of database performance optimization appeared first on Dynatrace news.

]]>
Data security graphic

With the recently released version of the Databases app, Dynatrace allows you to monitor your databases from a query perspective:

  • Quickly find all heavy queries that consume your database resources
  • Monitor your query metrics and resource utilization
  • Understand how your queries are executed through the Query Execution Plan

Optimize application performance: The importance of database analysis

Databases are critical for all (business-critical) applications and significantly affect performance. Monitoring the health and performance of your databases is essential to ensuring optimal application functionality. While infrastructure-level monitoring provides valuable insights, it might not reveal the root causes of database-related slowdowns. To address this, deeper analysis at the query level is necessary. By examining individual queries, you can identify database bottlenecks and implement targeted optimizations.

The Dynatrace Databases app introduces essential query-level monitoring features

The Databases app now offers long-awaited functionality for database monitoring at the query level:

Top Queries view lets you quickly identify queries that take a long time to execute or unnecessarily consume precious database resources. Additionally, it allows you to analyze how resource consumption or performance metrics for a given query have changed over time.

Query execution plans offer even more insights into how a database engine executes queries. Understanding what contributes most to query execution time is essential to query optimization. Instead of experimenting with different optimization techniques, deep analysis of the query execution plan makes pinpointing any possible resource conflicts or service degradations easy.

Top Queries and Execution Plans are available in the Databases app. These functionalities are currently supported for the following database engines:

  • Microsoft SQL Server
  • Oracle
  • PostgreSQL
  • MySQL
  • MariaDB

To start your analysis, open Databases, go to Instances > Top Queries and select the Statement Performance button for a specific database instance.

Databases app: Instances with Statement Performance
Figure 1. Databases app: Instances with Statement Performance

After identifying a query for analysis, you can look for interesting metrics that measure various conditions, like execution time, CPU consumption, or I/O utilization.

Statement performance analysis
Figure 2. Statement performance analysis

Expanding the query row allows you to analyze query performance over time from various perspectives.

Analyze queries across multiple perspectives.
Figure 3. Analyze queries across multiple perspectives.

To better understand how the given query is executed and to identify possible optimizations, you can request a query execution plan. Select the Request button and go to the Execution plan tab.

Query execution plan
Figure 4. Query execution plan

Execution plans provide a roadmap for how the database engine executes queries. With access to execution plans, you can identify performance bottlenecks, such as inefficient joins, missing indexes, and long-running queries, and optimize them for better efficiency and performance.

Execution plans also allow you to understand the cost and resource usage of the various query operations involved, enabling informed decisions about query re-design and indexing strategies. You can ensure that your databases run efficiently, and ultimately improve application performance.

Get started with the Databases app

Start monitoring your databases at the query level with the Dynatrace Databases app. Ensure that your databases are not a bottleneck for your apps and that you’re efficiently using database queries with the new and improved query execution plans.

If you have specific improvements in mind or would like to share feedback with us, please visit our Dynatrace Community feedback channel.

The post Analyze query performance: The next level of database performance optimization appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/analyze-query-performance-the-next-level-of-database-performance-optimization/feed/ 0
Easily troubleshoot z/OS application issues through logs with Dynatrace https://www.dynatrace.com/news/blog/troubleshoot-z-os-application-issues-logs-with-dynatrace/ https://www.dynatrace.com/news/blog/troubleshoot-z-os-application-issues-logs-with-dynatrace/#respond Wed, 09 Oct 2024 19:42:19 +0000 https://www.dynatrace.com/news/?p=66117 Explore logs graphics

New support for the IBM z/OS operating system automates log discovery and enables collection at scale. Get better hybrid cloud observability with the Dynatrace platform by including automatically enriched log data and faster issue troubleshooting.

The post Easily troubleshoot z/OS application issues through logs with Dynatrace appeared first on Dynatrace news.

]]>
Explore logs graphics

Logs become an integrated part of observability

Organizations face daily challenges in delivering integrated digital services essential to meeting their business goals. This is particularly true in hybrid cloud architectures, where system complexity, security, and performance are difficult to manage across cloud providers and the IBM Z mainframe platform.

A prerequisite for a successful hybrid cloud is maintaining observability, including all telemetry signals: logs, metrics, and traces. Without combining these signals in a unified AI-powered observability platform, the effectiveness of AIOps workflows in remediating problems is diminished, leading to wasted investment.

The Dynatrace® software intelligence platform can help you manage the complexity of digital services and hybrid clouds by providing holistic end-to-end visibility from the frontend, where your customers interact with your application, to the backend, where business transactions are processed.

Extend root cause analysis to logs on IBM z/OS

Dynatrace provides a platform for observing hybrid clouds and introduces support for log collection of the IBM z/OS operating system. This includes IBM CICS regions and IBM IMS subsystems.

You can now extend root cause analysis for any issue identified by Davis® AI with logs that are automatically linked to z/OS applications, transactions, or other identifiers specific to the environment hosting the resources.

Dynatrace offers log management and collection in a single place for public and private clouds and mainframe platforms such as IBM Z or LinuxONE. This makes log collection policies much more effective and transparent.

For example, you can apply a filter change to ignore certain logs from your central Dynatrace environment on all your monitored platforms without making any manual adjustments.

Centralized data masking rules allow administrators to easily configure the masking of sensitive data. This allows customers to address local or industry-specific regulatory or privacy requirements. The configured rulesets are directly deployed to the OneAgents, where the logs are collected.

Speed up your troubleshooting processes

Log analysis is typically one of the first steps in troubleshooting frontend problems. When a critical issue arises, it’s essential to have the right logs available to quickly and easily understand the full scope of what’s happening within your applications on the backend.

Dynatrace automatically discovers and collects logs from monitored IBM CICS regions and IBM IMS subsystems. All collected logs are enriched with metadata to map them to the entity model of z/OS hosts (logical partitions) and z/OS processes (regions and subsystems).

Dynatrace dramatically shortens Mean Time To Identify (MTTI) and Remediate (MTTR) times for incidents, as relevant log lines are provided in the context of detected problems.

Incidents are often not tied to a single component, as surrounding components and services can cause disruptions. With a single click, Dynatrace provides a surrounding logs view, showing the log lines of related components.

The newly released Dynatrace Logs app offers broad insights for manual investigation. Easy click-to-filter elements and a newly introduced DQL editor capable of translating these selected filters into Dynatrace Query Language (DQL) improve the experience for novice users.

Thanks to the enriched log data, log lines are connected to the respective z/OS Host pages.

Monitoring logs with Dynatrace facilitates novel ways to analyze telemetry data, significantly expanding the observability use cases for IBM Z mainframes. For example, with DQL queries, operators can quickly access all abends or drill down into specific job statistics.

Dynatrace Notebooks allows deep analysis by leveraging the Dynatrace Query Language within notebooks along with the newly introduced Davis CoPilot™ integration, a natural text-to-DQL builder.

Get started with logs on Dynatrace

Get started with Log Management and Analytics powered by Dynatrace Grail™ or Log Monitoring Classic (for Dynatrace Managed deployments). To start, deploy Dynatrace on your IBM z/OS operating system and set up log collection. You can easily set up log collection by turning on the provided z/OS log ingest rules globally or for a specific host.

Go to Settings > Log Monitoring > Log ingest rules and turn on z/OS CICS message user and z/OS IMS master terminal to start collecting logs.

  • Log monitoring on IBM z/OS is available with the release of Dynatrace OneAgent version 1.297 and ActiveGate version 1.297 (with the zRemote module) for Dynatrace SaaS and Dynatrace Managed.

The post Easily troubleshoot z/OS application issues through logs with Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/troubleshoot-z-os-application-issues-logs-with-dynatrace/feed/ 0
Simplicity meets power: Introducing the all-new Dynatrace Logs app https://www.dynatrace.com/news/blog/all-new-dynatrace-logs-app/ https://www.dynatrace.com/news/blog/all-new-dynatrace-logs-app/#respond Wed, 02 Oct 2024 16:35:32 +0000 https://www.dynatrace.com/news/?p=65907 Dynatrace Logs icon

As enterprises continue to optimize and consolidate various legacy IT solutions, managing and analyzing logs from existing log sources becomes more critical and complex.

The post Simplicity meets power: Introducing the all-new Dynatrace Logs app appeared first on Dynatrace news.

]]>
Dynatrace Logs icon

The new Dynatrace Logs app, fully powered by Grail™ data lakehouse, significantly enhances the experience for novice and seasoned users. Logs delivers unparalleled simplicity with powerful, sharable views and insights to address these needs.

The menu bar of the new Logs app provides simple click-to-filter options. You can select single log lines to surface more insights in the details pane and further click-to-filter options.
Figure 1. The menu bar of the new Logs app provides simple click-to-filter options. You can select single log lines to surface more insights in the details pane and further click-to-filter options.

Enhanced log ingestion and seamless out-of-the-box integration

The Logs app is supported by comprehensive log ingestion capabilities provided by the Dynatrace platform and Dynatrace Grail while interwoven with other use-case-specific Dynatrace Apps:

Hybrid
Cloud Native Logs in Dynatrace App Context
OneAgent Amazon Kinesis Data Firehose Clouds
OpenTelemetry  Amazon Log Forwarder Infrastructure & Operations
Syslog Azure Native Dynatrace Service integration Kubernetes
Log Ingest API Azure Log Forwarder Databases
Logstash Google Cloud logs Application Security

While this is just an overview of the commonly used ingest routes and methods you can leverage, the Dynatrace Hub offers the complete set, currently with 500 supported technologies, 150 additional public extensions, and 65 Dynatrace Apps.

Harness log source sprawl

It’s common for large enterprises to provide hundreds of applications to employees and to host more than a dozen critical line-of-business applications. “Digital workers are now demanding IT support to be more proactive,” is a quote from last year’s Gartner Survey

Understandably, a higher number of log sources and exponentially more log lines would overwhelm any DevOps, SRE, or Software Developer working with traditional log monitoring solutions. The Dynatrace logs in context and surrounding logs features ensure that Dynatrace users are provided with a view that’s tailored to their needs and supports their daily tasks.

The Logs app enables Dynatrace users, whether novice or experienced power users, to rapidly filter and investigate results manually within the app or experience a seamless drill-down handover from use-case-specific apps.

Logs are presented in the context of the applications that generate them, with the capability to run queries and open queried log entries directly in the Logs app.
Figure 2. Logs are presented in the context of the applications that generate them, with the capability to run queries and open queried log entries directly in the Logs app.

Whether it’s cloud applications, infrastructure, or even security events, this capability accelerates time to value by surfacing logs that provide the crucial context of what occurred just before an error line was logged. For example, if one of your customers unexpectedly uploaded a 1 GB file instead of a 1 MB file, was there an error with the buffer overflowing, or was the network stack unable to handle the unexpected load? Even more importantly, how was the error handled, and did the process end successfully for the customer? With Dynatrace and Surrounding logs, answers to such questions are never more than a click away, with no need to write complex queries or patterns.

You can filter surrounding logs from within the Logs app, or by an open-with call from another app (for example, the Kubernetes app shown in Figure 2).
Figure 3. You can filter surrounding logs from within the Logs app, or by an open-with call from another app (for example, the Kubernetes app shown in Figure 2).

Simplicity for novice and power users

For users who seek quick access to relevant logs without the need to write complex queries, easy filtering capabilities are available from within individual log line details or by adding and selecting fields in the menu bar.

For those who aspire to become power users, the new in-app DQL editor (Dynatrace Query Language) translates manually selected filters into the DQL code executed in the backend.

Logs app in advanced DQL-editor view, showcasing the capability to add additional filter conditions.
Figure 4. Logs app in advanced DQL-editor view, showcasing the capability to add additional filter conditions.

Power users can edit and enhance DQL queries to update results, or they can start directly within the DQL editor view. This allows deep-dive analysis when manual investigation or complex queries are required.

As the screenshot above shows, you can transition back to the filter selection menu bar by selecting Back to previous filters or by sharing the query and results with other teams.

This ensures a smooth user experience for DevOps engineers and SREs, whether they prefer intuitive click-and-filter workflows or fine-grained control through DQL.

Video of a Dynatrace user reviewing individual log lines, leveraging automatic log correlation of surrounding logs, and viewing details of additionally surfaced log entries.
Figure 5. Video of a Dynatrace user reviewing individual log lines, leveraging automatic log correlation of surrounding logs, and viewing details of additionally surfaced log entries.

Shortening time to value and increasing impact

With the general availability (GA) of the Logs app, another enhancement was introduced that shortens the time to value and increases the positive impact made within organizations.

Filtered log views that were customized by click-to-filter or by leveraging the DQL editor can now be shared using direct links—simply copy and paste the link displayed in your browser’s address bar.

This enables you to

  • bookmark and share reoccurring or complex queries.
  • share situation or incident-specific views across teams.
  • provision new Dynatrace users with relevant queries
  • share timeframe-specific and pre-filtered views in a support case

See more with less

During the preview release of the Logs app, many customers provided us with valuable feedback that we incorporated into the design of the app.

User requests, such as the newly added line wrap feature, increase the readability of results and eliminate the extra clicks required to open a detail view or load the results within a notebook. This shortens the time it takes to investigate results and take action.

Speaking of doing more with less, more was released simultaneously with the GA of the Logs app:

  • Increased performance of query results returning
  • Live search within query results
  • Live sorting within query results
  • custom column arrangement
  • enhanced open-with capabilities to open queries in other platform apps,
    or natively within the Logs app (for example, with Notebooks)
The open-with capabilities of the Logs app allow users to open single entries, such as individual host or node IDs in other apps, and to load the records in other applications like Notebooks or Dashboards.
Figure 6. The open-with capabilities of the Logs app allow users to open single entries, such as individual host or node IDs in other apps, and to load the records in other applications like Notebooks or Dashboards.

Why Dynatrace

The Dynatrace advantage lies in its ability to set logs in direct context, either automated by Davis® AI, in the context of specific team needs within an application, or as part of the broader observability scope provided by the Logs app itself.

Only Dynatrace provides a comprehensive and accessible log management and analytics experience, helping teams resolve issues faster without compromising on depth. Only Dynatrace Grail is schema-on-read and indexless, built with scaling in mind and built for exabyte scale, leveraging massively parallel processing. With Dynatrace, there is no need to think about schema and indexes, re-hydration, or hot/cold storage concepts.

This architecture also means you’re not required to determine your log data use cases beforehand or while analyzing logs within the new Logs app.

This includes logs collected from your organization’s Hypervisors, Linux and Windows servers, cloud-native logs from Azure, AWS, GCP, or Oracle, alongside the networking signals from Cisco, Juniper, NetScaler, or F5 appliances—to name a few examples.

What’s next

It’s simple, fast, and easy to ingest and gain multi-purpose value from logs—all without the need to be an expert or learn a complex query language first.

Learn how Dynatrace can address your specific needs with a custom live demo. Our observability experts will walk you through our solutions and show you how to deliver excellent customer experiences, foster your application security, and simplify IT operations.

If you’re not yet a Dynatrace customer, start your 15-day free trial.

If you want to learn more about Dynatrace and Logs in context, join us for a demo.

The post Simplicity meets power: Introducing the all-new Dynatrace Logs app appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/all-new-dynatrace-logs-app/feed/ 0
Unlock the power of contextual log analytics https://www.dynatrace.com/news/blog/unlock-the-power-of-contextual-log-analytics/ https://www.dynatrace.com/news/blog/unlock-the-power-of-contextual-log-analytics/#respond Wed, 02 Oct 2024 16:29:18 +0000 https://www.dynatrace.com/news/?p=65826 Observability graphic

In the ever-evolving landscape of IT operations and software development, logs are a critical data source for understanding system behavior, diagnosing issues, and maximizing business value. However, different teams often rely on a variety of monitoring and troubleshooting tools that use different data types, leading to fragmented data and inconsistent analytics. Dynatrace addresses this challenge by providing unified analytics and automation for logs, integrating them with all other observability, security, and business data types.

The post Unlock the power of contextual log analytics appeared first on Dynatrace news.

]]>
Observability graphic

Dynatrace enables various teams, such as developers, threat hunters, business analysts, and DevOps, to effortlessly consume advanced log insights within a single platform. Dynatrace Grail™ and Davis® AI act as the foundation, eliminating the need for manual log correlation or analysis while enabling you to take proactive action.

Dynatrace unified observability and security help enterprises, like our customer, BMO, save time and money, fostering collaboration across business, development, and operations teams.

In this blog post, we’ll provide an overview of the following log-related topics:

  • Logs in the context of applications
  • Easy yet powerful access to any log with the all-new Logs app
  • How logs are ingested
  • Dynatrace Application Security with logs
  • Dynatrace Davis CoPilot™ integration and AI-powered Application Security
  • Schemaless: Instantly gain business event insights from logs

Simplicity is key to success

As IT responsibilities shift left and expand, such as when developers take on application security duties, simplicity in toolsets becomes essential. Existing siloed tools lead to inefficient workflows, fragmented data, and increased troubleshooting times.

Tool consolidation is becoming a priority for C-level decision-makers in 2025. Enterprises are turning to Dynatrace for its unified observability approach for cloud-native, on-premises, and hybrid resources.

Rather than relying on disparate tools for each environment and team, Dynatrace integrates all data into one cohesive platform. Davis AI contextually aligns all relevant data points—such as logs, traces, and metrics—enabling teams to act quickly and accurately while still providing power users with the flexibility and depth they desire and need.

The Clouds app provides a view of all available cloud-native services. Logs in context, along with other details, are instantly available after selecting a resource.
Figure 1. The Clouds app provides a view of all available cloud-native services. Logs in context, along with other details, are instantly available after selecting a resource.

Logs in context with applications

Applications are provided within the Dynatrace platform to address the various needs of different teams and specific use cases. DevOps teams operating, maintaining, and troubleshooting Azure, AWS, GCP, or other cloud environments are provided with an app focused on their daily routines and tasks.

For instance, in a Kubernetes environment, if an application fails, logs in context not only highlight the error alongside corresponding log entries but also provide correlated logs from surrounding services and infrastructure components. This shortens root cause analysis dramatically, as explained in our recent blog post Full Kubernetes logging in context from Fluent Bit to Dynatrace.

The Kubernetes app provides an overview of the current log volume, criticality, and security status. Automatic log correlation for the selected Kubernetes node happens in the backend and is visualized when selecting Run query.
Figure 2. The Kubernetes app provides an overview of the current log volume, criticality, and security status. Automatic log correlation for the selected Kubernetes node happens in the backend and is visualized when selecting Run query.

The show surrounding logs function provides Dynatrace users with the ability to dive deeper and surface context-specific log lines of the components and services linked to the problem—all without a single line of code or complex query language knowledge. This is explained in detail in our blog post, Unlock log analytics: Seamless insights without writing queries.

For advanced analysis, there is a direct “open with” path, allowing you to load the current view in Notebooks for manual analysis, the Logs app, or other apps capable of visualizing context-specific log lines.

Screen video of Dynatrace platform when Davis AI automatically correlates Amazon AWS EC2 and business backend logs
Figure 3. A Service Reliability Engineer (SRE) manually reviews cloud-native front-end application warnings. Davis AI automatically correlates Amazon AWS EC2 and business backend logs. The platform offers the flexibility to dive deeper or filter views at any time by selecting highlighted components.

While the way that logs in context are interconnected and provided by Dynatrace is unique, as also Gartner® recognized for the 14th consecutive time in their Magic Quadrant™, there are use cases where this is not sufficient and raw access to the logs is required.

Easy yet powerful access to any log with the all-new Logs app

Developers love Dynatrace, without a doubt, and are one of the many teams next to DevOps or SREs making use of our all-new Logs app. The reasons are easy to find, looking at the latest improvements that went live along with the general availability of the Logs app.

In our product news blog post, Simplicity meets power: Introducing the all-new Dynatrace Logs app, we examine these features, which make life easy for new Dynatrace users, along with the newly introduced DQL Editor for power users.

­­Screenshot with unfiltered query results, where the newly introduced ’search in results’ feature has been used, to locally filter for log lines containing the word ’product‘.
­­Figure 4. Screenshot with unfiltered query results, where the newly introduced ’search in results’ feature has been used, to locally filter for log lines containing the word ’product‘.

Directly from individual log line results, you can filter simply by selecting corresponding items in the details pane or by loading the surrounding logs when selecting Show surrounding logs.

Keep in mind that Dynatrace Grail is schema-on-read and indexless, built with scaling in mind. There is no need to think about schema and indexes, re-hydration, or hot/cold storage. This architecture also means you are not required to determine your log data use cases beforehand or while analyzing logs within the new logs app.

How logs are ingested

Dynatrace offers OpenPipeline to ingest, process, and persist any data from any source at any scale. OpenPipeline ensures data security and privacy—data is collected and processed securely and compliantly, with high-performance filtering, masking, routing, and encryption—and contextualizes incoming data in real time. Using patent-pending high ingest stream-processing technologies, OpenPipeline currently optimizes data for Dynatrace analytics and AI at 0.5 Petabyte per day and tenant; this will soon increase to one Petabyte per day and tenant.

OpenPipeline’s high-performance filtering and preprocessing provide full ingest and storage control for the Dynatrace platform. As a result, dedicated data pipeline tools are unnecessary for preprocessing data before ingestion.

OpenPipeline architecture log flow
Figure 5. OpenPipeline architecture log flow

Dynatrace meets your teams where they are, with your preferred ingest routes and methods—be they Fluent Bit, OpenTelemetry, SysLog, or automatic log collection leveraging OneAgent, to name a few options.

If your team deploys applications cloud-natively, we meet you there, too, as we recently covered in our blog post, Dynatrace log management innovations: Syslog, AWS Firehose. We covered in rich detail how Dynatrace supports log ingestion for cloud-native workloads and simplified log ingestion also for hybrid environments.

In any case, at the heart of the Dynatrace Platform, Grail enables contextual analytics across unified observability, security, and business data. Grail is built for exabyte scale and leverages massively parallel processing (MPP) as well as advanced automated cold/hot data management to ensure that data remains fully accessible at all times, with zero latency, and full hydration.

Figure 6. Dynatrace marketecture – Logs in context
Figure 6. Dynatrace marketecture with logs in context

With no index or schema boundaries in place, paired with long-term data retention ranging from 1 day up to 10 years, you can leverage metrics, logs, and traces for a variety of additional use cases, such as business analytics or security analytics.

Dynatrace Application Security with logs

While most enterprises have Application Firewalls (AppFW), Intrusion Detection Solutions (IDS), and Static Code Analysis (SCA) in place for applications that will be deployed to production, it’s still relevant to understand if anomalies occur during runtime. Monitoring known vulnerabilities within the service hosting the application itself is just another puzzle piece to be considered for full end-to-end observability.

Instead of relying on static patterns, Dynatrace Causal AI understands the desired outcome of the triggered action and the context of the environment hosting the service. For example, deleting the database is not an expected outcome when the function provided is to update a user profile.

Dynatrace Security Investigator app visualizes the results of a shared incident investigation. The right-hand pane provides a query tree view with manually saved patterns in the evidence collection pane.
Figure 7. Dynatrace Security Investigator app visualizes the results of a shared incident investigation. The right-hand pane provides a query tree view with manually saved patterns in the evidence collection pane.

With these sophisticated security analytics, simple use cases such as authentication failure anomalies or password spraying attacks, along with more technical HTTP RST statistics, can be visualized in simple and sharable views in Security Investigator, leveraging logs.

Advanced analytics are not limited to use-case-specific apps. Dynatrace offers Notebooks and Dashboards to build views and reports—without the need to write a single line of code or Dynatrace Query Language, all supported by Dynatrace Davis CoPilot™ integration.

Dynatrace Davis CoPilot integration and AI-powered Application Security

Davis CoPilot™ assists occasionally visiting Dynatrace operators throughout the platform in a variety of applications, including Dynatrace Notebooks.

With natural language input in the example displayed in the screenshot below, “Show me the most recurring log lines and add a column with the log source and AWS region,” Davis CoPilot will evaluate the input, translate it into a corresponding DQL Query, and fetch the results accordingly.

Guardrails are in place and can be altered to prevent a high volume of unwanted material from being scanned or returned to Notebooks.

The Dynatrace operator defined a question in natural language that Davis CoPilot translated into Dynatrace Query Language
Figure 8. The Dynatrace operator defined a question in natural language that Davis CoPilot translated into Dynatrace Query Language

In the same way, Davis CoPilot can recommend remediation strategies and simplify security analysis across all data by translating natural language into the Dynatrace Query Language (DQL) to drive attack protection, security investigations, and forensics.

As mentioned, when ingesting relevant logs into Grail, including firewall and authentication gateway logs, Davis AI can provide end-to-end security insights. At the same time,

Davis AI not only visualizes or assesses risks automatically; it also detects and provides the option to block such threats when the corresponding features have been activated and configured platform-wide.
Figure 9. Davis AI not only visualizes or assesses risks automatically; it also detects and provides the option to block such threats when the corresponding features have been activated and configured platform-wide.

Davis CoPilot is integrated into the Dynatrace platform as an intelligent assistant to ease the effort of configuring and finding relevant details and options. By simply asking CoPilot the question stated above, you’re provided with the required configuration steps and product documentation references.

Davis CoPilot Assistant was asked to provide guidance for the example provided in this blog post and swiftly replied with the required configuration, referencing the source Application Security FAQ.
Figure 10. Davis CoPilot Assistant was asked to provide guidance for the example provided in this blog post and swiftly replied with the required configuration, referencing the source Application Security FAQ.

Schemaless: Instantly gain business event insights from logs

A unique picture can be drawn when logs from business-critical LoB applications are collected, which is an underestimated value that can be gained when using Dynatrace.

There are many customer examples and use cases, like room bookings in hotel portals, shopping cart statistics from online shops, or customer metrics from finance platforms, where customers can gain additional business value by translating logs to metrics within Dynatrace.

A pipeline health dashboard showing growth, paired with security-related information, is just one of the many examples of what you can build, either on your own or with the help of the Dynatrace Services team.

A custom business dashboard provides a holistic view of business process performance
Figure 11. A custom business dashboard provides a holistic view of business process performance

In our blog post, Leverage logs for an end-to-end view of your business processes via Dynatrace OpenPipeline, we demonstrated the retail company dashboard example above at a higher granularity, including the necessary steps for JSON log ingestion via Dynatrace OpenPipeline.

Conclusion

It’s simple, fast, and easy to ingest and gain multi-purpose value from logs—all without the need to be an expert or the requirement to first learn a complex query language.

Davis AI and CoPilot make it easy for casual contributors and power users to gain meaningful insights and build notebooks or dashboards while simultaneously increasing application security and reliability.

Learn how Dynatrace can address your specific needs with a custom live demo. Our observability experts will walk you through our solutions and show you how to deliver excellent customer experiences, foster your application security, and simplify IT operations.

If you’re not yet a Dynatrace customer, start your 15-day free trial.

If you want to learn more about Dynatrace and Logs in context, join us for a demo.

The post Unlock the power of contextual log analytics appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/unlock-the-power-of-contextual-log-analytics/feed/ 0
Next-level batch job monitoring and alerting: Elevate performance and reliability https://www.dynatrace.com/news/blog/next-level-batch-job-monitoring-and-alerting-elevate-performance-and-reliability/ https://www.dynatrace.com/news/blog/next-level-batch-job-monitoring-and-alerting-elevate-performance-and-reliability/#respond Fri, 27 Sep 2024 15:45:45 +0000 https://www.dynatrace.com/news/?p=65730 Dynatrace Security

Batch jobs are the backbone of automated, scheduled processes that execute tasks in bulk, such as data processing, system maintenance, or report generation. These jobs, which typically run in the background without user interaction, are critical and indispensable for handling large-scale operations efficiently.

The post Next-level batch job monitoring and alerting: Elevate performance and reliability appeared first on Dynatrace news.

]]>
Dynatrace Security

As batch jobs run without user interactions, failure or delays in processing them can result in disruptions to critical operations, missed deadlines, and an accumulation of unprocessed tasks, significantly impacting overall system efficiency and business outcomes. The urgency of monitoring these batch jobs can’t be overstated.

Monitor batch jobs

Monitoring is critical for batch jobs because it ensures that essential tasks, such as data processing and system maintenance, are completed on time and without errors. Failures, delays, or resource issues can lead to operational disruptions, financial losses, or compliance risks. Continuous monitoring enables early detection of problems, allowing quick remediation and maintaining business continuity.

Most jobs provide detailed information about job execution, including status, errors, and processing times in logs. The first step in monitoring batch jobs is to ingest these logs into Dynatrace. This is achieved by identifying the log files generated by the batch job program.

Apply basic filtering to ensure the availability of batch job-related logs. In this case, filter the logs based on relevant phrases or keywords.
Figure 1. Apply basic filtering to ensure the availability of batch job-related logs. In this case, filter the logs based on relevant phrases or keywords.

In this case, batch job statuses are constantly written from the deployment name get-cc-status-*. Thus we can create a rule in Dynatrace to ingest these logs via OneAgent without making any changes to the container, cluster, or host. Logs can also be ingested from various sources, including OpenTelemetry and Fluentbit.

A great reference is our blog post, Leverage edge IoT data with OpenTelemetry and Dynatrace, in which we documented the required steps to parse and ingest a single JSON log file into Dynatrace via OpenTelemetry.

Once logs are ingested, parsing the key messages is crucial. Below is a sample query that demonstrates how batch jobs can be parsed to extract important fields:

fetch logs
| filter matchesPhrase(content, "JOBS") AND matchesPhrase(content, "RunID")
| filter matchesValue(dt.entity.host, "HOST-HOSTID12345678")
| parse content , "
LD 'JOBS.' WORD:Job
LD 'RunID ' STRING:RunId
LD:status"
| fields timestamp, Job, status, content, RunId
| filterOut status == "."
| fieldsAdd start_time=if(contains(content,"started."),timestamp)
| fieldsAdd end_time=if(contains(content,"ended normally."),timestamp)
| fields timestamp, content, Job, status, RunId,start_time,end_time
Parsing the log lines that have critical data related to batch job status
Figure 2. Parsing the log lines that have critical data related to batch job status

Now that we can parse critical information, we can make informed decisions. However, it’s important to know if a job that started has ended within the expected timeframe. When a batch job exceeds its allotted time, the issue must be quickly identified and remediated.

Capture the time difference between two log entities

We use JavaScript within Dynatrace Dashboards to determine whether a previously started job was successfully completed. This three-level approach helps track how long a job took to complete and identifies any stuck jobs.

  1. Identify the unique property of each job and initialize its structure.
    const batch = {};
    
    /* Reiterate through each record and populate the data-structure*/
    for (const record of recordSet) {
      const runId = record['RunId'];
      if (!batch[runId]) {
        batch[runId] = {
          Job: record["Job"],
          run_id: runId,
          Status: "",
          JobStarted: null,
          JobEnded: null,
          Duration: "NA"
        };
      }
    }
  2. Process each job’s start time, end time, and status from the DQL parsed output.
    if (record["start_time"]) batch[runId].JobStarted = utcToLocal(record["start_time"]);
    if (record["end_time"]) batch[runId].JobEnded = utcToLocal(record["end_time"]);
    
    if (!statusLocked[runId]) {
      let status = record["status"]?.trim() || "";
    
      if (status.toLowerCase().includes("ended with return code")) {
        batch[runId].Status = "Failed";
        statusLocked[runId] = true;
      } else if (status == "started.") {
        batch[runId].Status = "Running";
      } else if (status == "ended normally.") {
        batch[runId].Status = "Completed without errors";
        statusLocked[runId] = true;
      } else {
        batch[runId].Status = status;
      }
    }
  3. Update the job status based on specific conditions (running, failed, completed).
    /* Leverage pre-populated data to identify duration for the completed jobs*/
    for (const runId in batch) {
      const job = batch[runId];
    
      if (job.JobStarted && job.JobEnded) {
        const startTime = new Date(job.JobStarted);
        const endTime = new Date(job.JobEnded);
        const duration = endTime - startTime;
    
        job.Duration = `${duration / 1000} seconds`;
      }
    }

Resources for the dashboard and workflow mentioned above can be found in this GitHub repository.

Individual batch job status with processing times and status
Figure 3. Individual batch job status with processing times and status
Advanced statistics for further analysis of batch jobs (median duration and job by status)
Figure 4. Advanced statistics for further analysis of batch jobs (median duration and job by status)

Correlate the impact of batch jobs with the application

Batch jobs should not impact applications because they run in the background. While they consume resources, they shouldn’t impact resource usage or client-facing applications. We can use Dynatrace Grail™ data lakehouse for unified observability data.

Correlate batch job runs with Application and Service resource utilization
Figure 5. Correlate batch job runs with Application and Service resource utilization

Adjust log parsing to account for varying log patterns

No two batch jobs are the same, and the log patterns you encounter might differ from what you see here. You can achieve the same results by parsing the logs. Parsing logs, as shown above, can be done using DPL Architect.

DPL Architect is a handy tool, accessible through the Notebooks app, that helps you quickly extract fields from records. It helps create patterns, provides instant feedback, and allows you to save and reuse DPL patterns, for faster access to data analytics use cases. This blog post offers further details about DPL architect.

Alerting for long-running or failed batch jobs

Constantly monitoring a dashboard isn’t practical, so you need automated alerting. Dynatrace workflows can check the status of batch jobs every 15 minutes and send alerts for failures or long-running jobs. These alerts can trigger actions or notifications sent via Slack, Teams, or as a ticket in your IT service management tool. In this example, the notifications are sent via email.

Automate batch job alerting and reporting
Figure 6. Automate batch job alerting and reporting

Conclusion

Monitoring batch jobs is essential to ensure they run smoothly and within expected timeframes. We can effectively identify issues such as long-running or failed jobs by ingesting logs into Dynatrace, parsing critical job information, and using custom logic to track job completion times. Implementing automated alerts triggering actions and notifications ensures proactive management, allowing teams to quickly resolve problems and maintain operational efficiency. With these tools in place, organizations can improve the reliability and performance of their batch-processing systems.

Use the approach detailed in this blog post to implement advanced batch job monitoring in your environment. Download the dashboards and Notebooks from this GitHub repository and start your automation journey today.

What’s next

In a future blog post, we’ll show how batch job management can be efficiently orchestrated using workflows and predictive analysis to schedule and run jobs optimally. With Davis® AI identifying root causes, workflows can be used to stop erroneous batch executions.

Additionally, Davis® AI prediction analysis, in conjunction with workflows, can reschedule or pause jobs to ensure optimal resource utilization, preventing any negative impact on the application landscape.

Download the Dashboards and Notebooks used in this blog post from our GitHub repository.

The post Next-level batch job monitoring and alerting: Elevate performance and reliability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/next-level-batch-job-monitoring-and-alerting-elevate-performance-and-reliability/feed/ 0
From syslog to AWS Firehose: Dynatrace log management innovations that enhance observability https://www.dynatrace.com/news/blog/from-syslog-to-aws-firehose-dynatrace-log-management-innovations-that-enhance-observability/ https://www.dynatrace.com/news/blog/from-syslog-to-aws-firehose-dynatrace-log-management-innovations-that-enhance-observability/#respond Thu, 05 Sep 2024 13:01:14 +0000 https://www.dynatrace.com/news/?p=65416 log management innovations

That first mile of getting data in can often be the hardest. That's why Dynatrace continues to invest in log ingest, offering a range of out-of-the-box solutions. With these latest log management innovations, you can harness even more data for comprehensive AI-driven insights, faster troubleshooting, and improved operational efficiency whether you use Syslog, AWS Firehose, Fluent Bit, or other technologies.

The post From syslog to AWS Firehose: Dynatrace log management innovations that enhance observability appeared first on Dynatrace news.

]]>
log management innovations

Understanding that the first mile of getting data in can often be the hardest, Dynatrace continues to invest in log ingest, offering a range of out-of-the-box solutions within the Dynatrace Platform and apps. We’re excited to announce several log management innovations, including native support for Syslog messages, seamless integration with AWS Firehose, an agentless approach using Kubernetes Platform Monitoring solution with Fluent Bit, a new out-of-the-box ingest dashboard, and OpenPipeline ingest improvements.

These developments open up new use cases, allowing Dynatrace customers to harness even more data for comprehensive AI-driven insights, faster troubleshooting, and improved operational efficiency.

Let’s delve deeper into how these capabilities can transform your observability strategy, starting with our new syslog support.

Native support for Syslog messages

Syslog messages are generated by default in Linux and Unix operating systems, security devices, network devices, and applications such as web servers and databases. Native support for syslog messages extends our infrastructure log support to all Linux/Unix systems and network devices. The more data ingestion channels you provide to the Dynatrace Davis® AI engine, the more comprehensive Dynatrace automated root cause analysis becomes.

Customers can also proactively address issues using Davis AI’s predictive analytics capabilities by analyzing network log content, such as retries or anomalies in performance response times.

Dynatrace natively supports Syslog using ActiveGate (preferred method) or the OpenTelemetry (OTel) collector. Many syslog producers lack authentication and have varying security capabilities, such as TLS encryption. Dynatrace ActiveGate addresses these issues by enforcing configurable security settings and ensuring data uniformity. It also enhances syslog messages with additional context and optimizes network traffic, improving overall system resilience and security. This enhancement lays the foundation for the broader, more integrated observability that Dynatrace continues to expand with each new feature.

Syslog endpoint ingest with Dynatrace diagram

Customers have had a positive response to our native syslog implementation, noting its easy setup and efficiency. A $20 billion Germany-based financial services company told us they found the process of pushing Syslog messages to Dynatrace natively to be seamless. Another customer based in Germany, a $23 billion medical technology company, told us they appreciate the value of using a native channel to push syslog messages from network devices directly to Dynatrace, bypassing the need for FluentD or a standalone OpenTelemetry collector.

This streamlined approach enhances both usability and integration, making syslog management simpler and more effective.

Seamless integration with AWS Firehose

Dynatrace is also enhancing our observability logs offerings for AWS services for cloud-native applications. By integrating AWS Firehose into the Dynatrace platform, you can address high-impact issues quickly through real-time, high-frequency log analytics.

This integration with AWS Firehose simplifies observability by removing intermediary components, which allows seamless log capture and analysis directly in the Grail data lakehouse. Logs are immediately available for troubleshooting, security investigations, and auditing, becoming integral to the platform alongside traces and metrics.

Dynatrace supports scalable data ingestion, ensuring your observability infrastructure grows with your cloud environment. The setup is straightforward, using API keys, CloudFormation templates, or the AWS web console.

Dynatrace also provides contextual insights by linking logs to problems detected by Davis AI, enabling quick access to relevant details. The platform also offers proactive analysis through Notebooks for visualizing log data and exploring error rates. Dynatrace support for AWS Firehose includes Lambda logs, Amazon virtual private cloud (VPC) flow logs, S3 logs, and CloudWatch. This seamless integration not only enhances AWS observability but also ties into the greater context of how Dynatrace unifies cloud-native observability across multiple platforms.

Customers have responded enthusiastically to our AWS Firehose implementation. Our approach provides seamless cloud log integrations you can configure directly in the AWS console or through provided CloudFormation templates. This setup eliminates the need for additional middle-layer components, making it straightforward and efficient.

A key advantage of this integration is its high throughput aligned with Grail, ensuring optimal performance. What’s more, logs ingested using AWS Firehose are enriched with cloud context, enabling in-context analysis within the platform. This capability enhances the overall observability and insights that customers can gain from their cloud environments.

log management innovations include support for AWS Firehose

Kubernetes Platform Monitoring using Fluent Bit for cloud-native environments

One of the log management innovations we’re excited to share is the new Kubernetes platform monitoring solution with Fluent Bit, offering a cloud-native, API-based deployment model. With this innovation, Dynatrace makes it easier for teams to stream logs from Kubernetes environments into Dynatrace through a more lightweight and streamlined setup, accelerating time to value. Customers get advanced health analytics out of the box and automated root cause analysis by Davis® AI when they ingest Kubernetes workloads, traces, logs, and metrics into the Dynatrace Grail data lakehouse.

For organizations who already use Fluent Bit as part of their tech stack to configure pipelines and enrich log data, this modern approach enables teams to gain answers in context based on logs, powered by the Dynatrace platform’s automation and problem detection. With all data in one place and in context, the new Kubernetes platform monitoring solution provides easy filters by namespace, cluster, workloads, nodes, services, pods, and containers. Integrating with Fluent Bit for Kubernetes log ingestion is important for ensuring teams are capturing critical data for troubleshooting and issue remediation.

Dynatrace enhances Fluent Bit’s log management by integrating observability signals like traces, events, and metrics, providing a complete view of cloud-native application performance. It automates log analysis, eliminates manual correlation, and offers broader visibility through ready-made dashboards and health checks. Log configuration is simplified, while advanced analytics powered by Davis AI bubbles up critical health signals, and provides automated root cause analysis, predictive AI for remediation, generative AI for query writing, performance baselining, and anomaly detection. This innovation ties into the broader effort to simplify and enhance log management across diverse environments, further integrating Kubernetes observability into the Dynatrace ecosystem.

Out-of-the-box logs ingest dashboard

Coming soon is an out-of-the-box logs ingest dashboard that enables you to easily manage your ingest channels.

The dashboard tracks a histogram chart of total storage utilized with logs daily. It also tracks the top five log producers by entity. You can see in a table retention periods by the number of logs and storage they consumed.

The dashboard also breaks down log volume by Grail buckets, showing you what buckets consume the most storage. Grail buckets can help enterprises categorize their logs by retention periods, types of logs such as audit logs, or by organizations that utilize the logs. Think of it like individual bookshelves in a library, where each shelf is dedicated to a specific genre or topic. Just as books are organized on these shelves for easy access and retrieval, data log records are stored in buckets based on their type or purpose, allowing for efficient management and quick querying. This organization ensures that when you need specific information, you can go directly to the relevant shelf or bucket, saving time and effort in finding what you need.

This final piece of the puzzle ensures that your log data is not only easily ingested but also effectively managed, tying together the full spectrum of observability enhancements in the Dynatrace platform. The logs ingest dashboard is currently in tech preview, with GA expected soon. Please follow up with your account team to get early access.

Ingest dashboard in Dynatrace screenshot

OpenPipeline ingest improvements save money and improve query performance

Lastly, the Dynatrace feature OpenPipeline unifies how we ingest, transform, enrich, and process all observability signals, including logs. This is a significant upgrade to our log processing pipeline capabilities. We now support the following for log ingest:

  • Log content up to 512K each
  • Log metric counts up to 1000
  • Log attributes up to 2.5KB
  • Number of log attributes up to 250
  • Logs ingestion API payload up to 10MB
  • Converting logs into Business Events saves money
  • Masking sensitive data
  • Setting security context
  • Extracting metrics from logs and business events saves money
  • Parsing JSON before storing logs in Grail for faster analytics

Customers can save money by converting logs into metrics and business events, which are ingested into predefined buckets, making queries faster without needing to span the entire Grail dataset. This approach also enhances security by allowing customizable security contexts for each log and masking sensitive data before ingestion, while simplifying analytics by pre-parsing JSON. With OpenPipeline, customers can add, remove, rename fields, parse, and mask all incoming logs.

Building on these ingest improvements, these innovations further enhance data analytics and precision by introducing advanced features like OpenPipeline.

Dynatrace log management innovations expand data analytics and precision

Dynatrace continues to lead the way in log management and observability with its latest advancements. By introducing support for syslog messages, AWS Firehose integration, the agentless Fluent Bit setup for Kubernetes environments, a powerful out-of-the-box ingest dashboard and our new OpenPipeline, our customers can achieve deeper insights, faster troubleshooting, and more efficient operations across their hybrid cloud ecosystems. These innovations not only expand the breadth of data available for analysis within the Dynatrace platform but also enhance the precision and effectiveness of our AI-driven capabilities, which results in more actionable, automatable outcomes.

Each of these developments interconnects to create a more comprehensive and powerful observability solution, designed to meet the evolving needs of our customers. As we continue to refine and expand our offerings, our commitment remains focused on empowering organizations to proactively manage their infrastructure and applications with unparalleled clarity and confidence.

We encourage our customers to explore these new capabilities and see how they can further optimize their observability practices.

Try out these new Dynatrace log management innovations

If you’re a customer, go to Dynatrace Playground to check out the new capabilities. If you’re looking into Dynatrace, check out our free trial.

To learn more about these technologies, see details in the following blogs:

State of Log Management 2026

Download the report to explore benchmark data on how AI workloads are exploding log volume and costs, and why unified observability is now essential for reliable, trustworthy AI.

The post From syslog to AWS Firehose: Dynatrace log management innovations that enhance observability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/from-syslog-to-aws-firehose-dynatrace-log-management-innovations-that-enhance-observability/feed/ 0
Six causes of major software outages–And how to avoid them https://www.dynatrace.com/news/blog/six-causes-of-major-software-outages-and-how-to-avoid-them/ https://www.dynatrace.com/news/blog/six-causes-of-major-software-outages-and-how-to-avoid-them/#respond Thu, 08 Aug 2024 14:00:16 +0000 https://www.dynatrace.com/news/?p=65055 software outages

Avoiding major software outages is an essential goal of business resilience plans for any industry. This blog is part of a series that explores how organizations can maintain business resilience to avoid—and recover from—IT outages.

The post Six causes of major software outages–And how to avoid them appeared first on Dynatrace news.

]]>
software outages

As recent events have demonstrated, major software outages are an ever-present threat in our increasingly digital world. From business operations to personal communication, the reliance on software and cloud infrastructure is only increasing.

Outages can disrupt services, cause financial losses, and damage brand reputations. Understanding the causes of these outages is crucial for preventing them and ensuring smoother, more reliable tech operations. It’s also critical to have a strategy in place to address these outages, including both documented remediation processes and an observability platform to help you proactively identify and resolve issues to minimize customer and business impact.

How software outages happen

Outages can occur for many reasons, ranging from internal mishaps to external attacks. They may stem from software bugs, cyberattacks, surges in demand, issues with backup processes, network problems, or human errors. Each of these factors can independently cause a major disruption, but often, outages result from a combination of issues. Let’s explore each of these elements and what organizations can do to avoid them.

1. Software bugs

Software bugs and bad code releases are common culprits behind tech outages. These issues can arise from errors in the code, insufficient testing, or unforeseen interactions among software components.

Possible scenarios

  • A new software update contains a bug that causes a critical application to crash, disrupting business operations.
  • A poorly tested feature release leads to incompatibility issues, resulting in downtime for users.

To prevent outages caused by software bugs, organizations should implement thorough testing procedures, including automated testing and continuous integration practices. Regular code reviews and a robust quality assurance process are also vital to help identify issues before they reach production.

2. Cyberattack

Cyberattacks involve malicious activities aimed at disrupting services, stealing data, or causing damage. These attacks can be orchestrated by hackers, cybercriminals, or even state actors.

Possible scenarios

  • A Distributed Denial of Service (DDoS) attack overwhelms servers with traffic, making a website or service unavailable.
  • Ransomware encrypts essential data, locking users out of systems and halting operations until a ransom is paid.
  • Remote code execution (RCE) vulnerabilities, such as the Log4Shell incident in 2021, allow attackers to run malicious code on a remote system without requiring authentication or user interaction.

To cope with the risk of cyberattacks, companies should implement robust security measures combining proactive preventive measures such as runtime vulnerability analytics, with comprehensive application and perimeter protection through firewalls, intrusion detection systems, and regular security audits. Employee training in cybersecurity best practices and maintaining up-to-date software and systems are also crucial.

3. High demand

Sudden spikes in demand can overwhelm systems that are not designed to handle such loads, leading to outages. This often occurs during major events, promotions, or unexpected surges in usage.

Possible scenarios

  • A retail website crashes during a major sale event due to a surge in traffic.
  • An online streaming service experiences downtime during the premiere of a highly anticipated show, as too many users try to access it simultaneously.

To manage high demand, companies should invest in scalable infrastructure, load-balancing, and load-scaling technologies. Conducting performance testing and having contingency plans for peak times can help ensure systems remain operational during spikes in usage.

4. Backup process

Failures in the backup process can lead to outages, especially when primary systems fail, and backups do not activate as expected. This can result from improperly configured backups, corrupted data, or insufficient testing.

Possible scenarios

  • A data center experiences a power failure, but the backup generators fail to start, leading to prolonged downtime.
  • A company tries to restore a system from backups after a cyberattack, only to find the backups are corrupted or incomplete.

It’s critical to regularly perform backup and recovery tests to ensure that systems are properly configured. Companies should ensure they have a range of recovery options in place, including snapshots, replication, and backups to provide a range of RTO and RPO options. A comprehensive DR plan with consistent testing is also critical to ensure that large recoveries work as expected.

5. Network issues

Network issues encompass problems with internet service providers, routers, or other networking equipment. These can be caused by hardware failures, or configuration errors, or external factors like cable cuts.

Possible scenarios

  • A major network provider experiences an outage, causing disruptions to services that rely on its infrastructure.
  • Misconfigured network settings result in lost connectivity, impacting cloud services and online applications.

To mitigate network issues, organizations should ensure robust network monitoring and management practices. Redundant network paths and automated failover systems can help maintain connectivity during disruptions.

6. Human error

Human error remains one of the leading causes of tech outages. This can include mistakes made during routine maintenance, misconfigurations, or accidental deletions.

Possible scenarios

  • An IT technician accidentally deletes a critical database, causing a service outage.
  • Incorrectly applied configuration changes lead to system failures and downtime.

Comprehensive training programs and strict change management protocols can help reduce human errors. Automated systems for routine tasks and thorough review processes for critical actions can also minimize the risk of mistakes.

Mitigating the causes of software outages

Understanding the diverse causes of tech outages is essential for developing strategies to prevent them, but it’s just the start. An effective mitigation strategy requires an observability solution that provides a complete end-to-end view of all applications and services. A platform such as Dynatrace enables companies to proactively identify issues, prioritize remediation, and validate that implemented fixes address the underlying issues. This approach minimizes the impact of outages on end users and maximizes the efficiency of IT remediation efforts.

The unfortunate reality is that software outages are common. However, by understanding the root causes of outages and implementing an observability platform, organizations can enhance the reliability and resilience of their technology infrastructure, ensuring continuity and maintaining trust in an increasingly digital world.

Contact us to learn how you can mitigate the causes of software outages in your IT environment to maintain business resilience.

To learn more about the recent CrowdStrike update outage and explore more resources to help you maintain business resilience, check out the resource center, Business Resilience through CrowdStrike and Beyond.

The post Six causes of major software outages–And how to avoid them appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/six-causes-of-major-software-outages-and-how-to-avoid-them/feed/ 0
CrowdStrike update: How Dynatrace helped customers recover in hours https://www.dynatrace.com/news/blog/crowdstrike-update-crisis-dynatrace-customers-recovery/ https://www.dynatrace.com/news/blog/crowdstrike-update-crisis-dynatrace-customers-recovery/#respond Wed, 31 Jul 2024 13:29:47 +0000 https://www.dynatrace.com/news/?p=65026 CrowdStrike update

The ability to recover quickly in case of a sudden IT outage is crucial for business resilience. This blog is part of a series that explores how organizations can maintain business resilience by having the right capabilities to recover quickly from an IT outage.

The post CrowdStrike update: How Dynatrace helped customers recover in hours appeared first on Dynatrace news.

]]>
CrowdStrike update

On July 19th, 2024, countless organizations had their operations disrupted by a routine software update from CrowdStrike, a popular cybersecurity software. The resulting outages wreaked havoc on customer experiences and left IT professionals scrambling to quickly find and repair affected systems.

A wide variety of companies and industries have suffered the effects of this incident, from delayed flights to disruptions in healthcare, insurance, and the financial industry. The ripple effects on the global supply chain have been equally significant. The crisis has emphasized the importance of having a strategy for maintaining stability and performance.

Time is of the essence in any crisis—so is having the right tools and capabilities. Although Dynatrace can’t help with the manual remediation process itself, end-to-end observability, AI-driven analytics, and key Dynatrace features proved crucial for many of our customers’ remediation efforts.

Following are some of the critical capabilities many of our customers relied on to recover from the CrowdStrike update crisis in just hours.

1. Real-time monitoring with out-of-the-box features

Real-time data and monitoring are crucial for maintaining situational awareness of IT environment stability and performance, especially during a crisis. Knowing what’s offline and which dependencies connect to mission-critical services is key to determining the impact of an incident and determining where to start remediation. Dynatrace offers various out-of-the-box features and applications to provide a high-density overview of system health for all hosts and related metrics in a single view.

Smartscape topology mapping

Dynatrace Smartscape® provides a dynamic, real-time visualization of the entire application topology so teams can quickly identify and address issues. Understanding application dependencies helps teams prioritize what to address first. For example, a good course of action is knowing which impacted servers run mission-critical services and remediating those first.

Problems application

The Problems application automatically identifies issues, collects the context behind them, and presents their root cause and impacts in a single view. Powered by Davis® AI, the app helps teams immediately see the problem’s duration, root cause, and business impact.

Dashboards and visualizations

Standard dashboards and visualizations also provide situational awareness out of the box. The following honeycomb visualization shows a healthy environment before an incident and after the incident starts. All the problems, offline hosts, databases, and failing services appear in red.

Honeycomb visualization: Before CrowdStrike outage
Systems before the outage
Honeycomb visualization: CrowdStrike outage in progress
Systems during the outage

Dynatrace OneAgent full-stack and Foundation and Discovery modes

Dynatrace OneAgent provides automatic discovery and monitoring. In addition to using OneAgent for full stack monitoring of the most critical applications, Dynatrace offers OneAgent Foundation and Discovery mode. This lightweight alternative provides full coverage of an environment in scenarios where teams need cost-effective yet comprehensive monitoring. Foundation and Discovery provide essential metrics and topology discovery, making it useful to quickly identify and recover affected hosts.

Together, these technologies enable organizations to maintain real-time visibility and control, swiftly mitigating the impact of incidents and efficiently restoring critical services. They also enable companies to measure the effectiveness of their remediation activities to ensure that recoveries proceed as expected.

How out-of-the-box monitoring features helped one company recover from the CrowdStrike outage within hours

A US-based pharmaceutical company recovered its most critical systems within hours of the CrowdStrike incident using out-of-the-box real-time monitoring features. Dynatrace automatically found the hosts that were unavailable or having problems. The key information displayed on the standard Dynatrace Problems app and the Infrastructure and Operations App became the basis of their team’s remediation plan. The company was back to normal business operations as other companies continued to struggle with recovering days after the initial software push.

2. Synthetic monitoring

Synthetic monitoring is a critical tool for ensuring application reliability and performance, especially during a crisis. By simulating user interactions and running tests from various locations worldwide, synthetic monitoring provides a comprehensive view of application performance and availability. This proactive approach allows organizations to detect and resolve issues early, optimize performance, and maintain a high-quality user experience.

Organizations can use synthetic monitoring to continuously monitor API endpoints and ensure that critical user journeys perform well and meet service level agreements (SLAs). This awareness can help reduce downtime and minimize disruption, enabling swift action if an incident occurs.

Many businesses rely on third-party services, such as payment processors, content delivery networks (CDNs), and ticketing systems to get through their day-to-day operations. Even if the business isn’t directly affected by a crisis, their third-party suppliers may be, which can still disrupt their operations.

How synthetic monitoring helped one company get an early warning of the CrowdStrike impact

During the recent CrowdStrike crisis, a US life insurance provider was impacted indirectly through its third-party ticketing system. The Dynatrace synthetic monitors they used to monitor the performance of this critical third-party application immediately detected the outage when the CrowdStrike software affected the vendor’s servers.

Dynatrace created a problem notification and Davis AI determined the root cause on the vendor’s side, hours before the company publicly announced the CrowdStrike incident affected it. This advanced warning allowed the life insurance provider to execute a contingency plan and course of action much earlier than if they had to wait for the third-party provider to notify them of the problem.

Synthetic monitoring view of the CrowdStrike outage

Synthetic monitoring view of the CrowdStrike outage
Synthetic monitoring views of the CrowdStrike outage

3. Dynatrace Query Language (DQL)

Dynatrace Query Language (DQL) is a structured syntax for exploring, querying, and processing observability data in Dynatrace. It allows users to chain commands together to filter, manipulate, and analyze data efficiently.

As a data query tool, DQL provides flexibility and customization, allowing organizations to tailor their investigations to meet specific needs, addressing unique challenges and optimizing performance.

Using DQL, investigators can find specific answers so they can quickly identify and respond to issues, which is crucial during crises like the CrowdStrike incident.

Dynatrace Notebooks is an interactive capability that enables users across the organization to collaborate using code, text, and rich media to build, evaluate, and share insights for exploratory analytics. This ability to track and collaborate on issue details is a crucial capability in a crisis.

How DQL and Notebooks helped companies pinpoint affected systems and prioritize remediation

During the CrowdStrike crisis, a North American telecommunications provider used DQL and a notebook to create custom charts filtered by application. This helped the company prioritize remediating its most critical servers first, restoring essential services promptly.

DQL and Notebooks investigation showing systems affected by the CrowdStrike outage
DQL and Notebooks investigation showing systems affected by the CrowdStrike outage

When a major US airline began experiencing the CrowdStrike outage, its IT team also used DQL and Dynatrace Notebooks to identify which systems were no longer forwarding logs back to Dynatrace. This newfound visibility enabled the kiosk management team to focus their recovery efforts effectively.

To see an example of how to use DQL to find when BSOD issues are being written to Windows system logs, see the blog Crowdstrike BSOD: Quickly find machines impacted by the CrowdStrike issue by Dynatrace Principal Solutions Engineer Josh Wood, Ph.D.

4. Real user monitoring to understand business impact

Real User Monitoring (RUM) offers comprehensive insights into user experiences across web, mobile, and custom applications. By capturing user sessions, RUM provides a detailed view of user journeys, helping businesses understand critical actions for conversions.

When an incident occurs, Dynatrace automatically generates a problem notification. Davis AI analyzes details from the front end to the backend to identify the root cause, severity, and impact of the issue.

Dynatrace RUM also tracks application downtime, enabling organizations to calculate the cost of business interruptions and account for lost revenue. RUM offers flexible options for tracking and reporting key information, such as conversion goals. Examples include successful checkouts, newsletter signups, or demo requests. By monitoring conversion rates, businesses can estimate expected revenue to better understand the financial impact of an incident and help organizations account for lost revenue.

How Dynatrace RUM helped one company identify the business impacts of the CrowdStrike outage

For example, during the CrowdStrike crisis, a North American mortgage provider received an alert for unexpected low traffic. The problem card helped them identify the affected application and actions, as well as the expected traffic during that period. This allowed them to prioritize remediation efforts on their most critical services.

Dynatrace RUM shows the user impact of the CrowdStrike outage
Dynatrace RUM shows the user impact of the CrowdStrike outage

5. Using SLOs to verify recovery

Service level objectives (SLOs) are essential for maintaining and enhancing the performance of applications and services, especially during and after a crisis. Dynatrace makes it easy to create, capture, and visualize SLOs in real time. Establishing and monitoring SLOs can play an instrumental role before, during, and after a crisis.

  • Before a crisis. Setting up SLOs for mission-critical services helps establish and maintain standards for availability and performance. Dynatrace AI continuously monitors these benchmarks, allowing teams to identify and address potential issues proactively.
  • During a crisis. SLOs provide real-time monitoring and immediate feedback on service performance. Dynatrace AI can quickly pinpoint the root cause of issues, enabling swift resolution and minimizing user impact.
  • After a crisis. SLOs ensure that application performance returns to the same standard of performance as before the incident. They play a crucial role in post-incident analysis, helping teams understand the business impact of the incident and implement improvements to prevent future occurrences.

By implementing Dynatrace SLOs, organizations can ensure robust performance management before, during, and after a crisis, leading to more resilient and reliable services.

Prepare for any crisis with observability and the right capabilities

The recent CrowdStrike crisis has highlighted the critical need for robust monitoring and observability tools. For organizations navigating disruptions, observability is crucial for rapid detection and remediation. Dynatrace offers comprehensive solutions with real-time data, synthetic monitoring, DQL querying capabilities, real user monitoring, and SLOs. These tools empower organizations to maintain stability, swiftly identify and resolve issues, and ensure the continuity of essential services.

The real-world examples of our customers demonstrate how Dynatrace monitoring solutions have enabled organizations to recover quickly and maintain operational stability, ultimately safeguarding their customer experiences and bottom lines. As we continue to lean heavily on technology for day-to-day operations, observability will be essential for navigating the complexities of an increasingly digital world.

Using Dynatrace, organizations can not only react fast to mitigate the immediate impacts of crises but also be proactive and build resilient IT infrastructure that’s prepared for future challenges.

Contact us to learn how you can gain the same situational awareness and responsiveness that enabled these customers to recover so quickly from the CrowdStrike outage.

To learn more about the recent CrowdStrike update outage and explore more resources to help you maintain business resilience, check out the resource center, Business Resilience through CrowdStrike and Beyond.

The post CrowdStrike update: How Dynatrace helped customers recover in hours appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/crowdstrike-update-crisis-dynatrace-customers-recovery/feed/ 0
A Kubernetes platform engineering strategy tames Kubernetes complexity https://www.dynatrace.com/news/blog/kubernetes-platform-engineering-tames-kubernetes-complexity/ https://www.dynatrace.com/news/blog/kubernetes-platform-engineering-tames-kubernetes-complexity/#respond Thu, 25 Jul 2024 07:51:02 +0000 https://www.dynatrace.com/news/?p=64943 Kubernetes platform engineering tames Kubernetes complexity; Dynatrace Kubernetes monitoring

With Kubernetes, it’s easy for organizations to miss the forest for the trees. Despite its becoming an industry standard, Kubernetes complexity can cause major headaches as organizations scale. This complexity prompted digital bank PicPay, Latin America’s largest digital wallet and now bank, to adopt a Kubernetes platform engineering approach with unified observability and security at its center.

The post A Kubernetes platform engineering strategy tames Kubernetes complexity appeared first on Dynatrace news.

]]>
Kubernetes platform engineering tames Kubernetes complexity; Dynatrace Kubernetes monitoring

After a decade of helping companies manage container orchestration, Kubernetes, the open source container platform, has established itself as a mature enterprise technology. According to the Cloud Native Computing Foundation (CNCF), 84% of organizations are using or evaluating Kubernetes, up from 81% in 2022.

But challenges remain when it comes to Kubernetes complexity. The average deployment now spans 20 clusters running 10 or more software elements across clouds and data centers. In fact, 76% of technology leaders say the dynamic nature of Kubernetes makes it more difficult to maintain visibility of their infrastructure compared with traditional technology stacks.

With Kubernetes, it’s easy for many organizations to miss the forest for the trees. While its open source attributes have helped establish it as the industry standard for deploying, managing, and scaling container-based applications, its complexity rapidly increases as Kubernetes deployments begin to scale.

This was the case for PicPay, a financial services app in Brazil. I spoke with Martin Spier, PicPay’s VP of Engineering, about the challenges PicPay experienced and the Kubernetes platform engineering strategy his team adopted in response.

Overcoming Kubernetes complexity

With more than 35 million active users, PicPay is experiencing substantial growth: In 2023, the Total Payment Volume jumped 40% to R$271 billion, while revenue reached a record of R$3.5 billion.

The company receives tens of thousands of requests per second on its edge layer and sees hundreds of millions of events per hour on its analytics layer. To manage this data, PicPay has more than a hundred Kubernetes clusters running tens of thousands of pods and over a thousand microservices. “We ended up with allegedly the largest cluster in Latin America” Spier said, “which isn’t a great thing”.

Moreover, observability of their increasingly complex Kubernetes environment was lagging. “Our development teams relied heavily on logs to understand what was going on with our systems,” he said. This created problems with both visibility and scalability. Different teams were using different solutions to achieve the same results, creating a fragmented IT stack. In addition, their logs-heavy approach to analysis made scaling processes complex and costly. By over-rotating on log analysis, Spier and his team were missing the value, cost savings, and productivity that come from having metrics, traces and logs all in one place and in context.

To address these problems, Spier and his team needed to simplify. “We decided to break up the big cluster into smaller ones and create a standardization to provide that managed infrastructure for everyone,” Spier said. The company’s goal was to standardize observability and prevent common problems, such as Java or pods running out of memory, or users requesting resources and barely using any, or using 100% of it.

Taking a strategic Kubernetes platform engineering approach

Spier noted that keeping Kubernetes simple requires a strategic approach. He points to the shift from DevOps to platform engineering, or as he calls it, Foundation Engineering.
“Software is built in layers,” Spier explained. “And these layers tend to be similar. But if every team is left to define their own tools and stacks, you end up duplicating a lot of the work. Platform engineering looks to bring in a unified toolset.”

Spier also noted that trying to support all possible use cases is another common pitfall in navigating Kubernetes complexity. “For example, if most teams run Java, it might not make sense trying to support an outlier. Instead, you’re better off creating the best solution for the common use case.”

Ultimately, the sooner companies start controlling complexity using a standardized Kubernetes platform engineering approach, the better. “Start before you have multiple, competing platforms and you have to go through a really unpleasant migration,” Spier suggested.

Unified observability is a team sport for taming Kubernetes complexity

To implement their Kubernetes platform engineering strategy, PicPay needed observability of the big picture. They also needed to integrate the value and context of metrics and traces into their log monitoring scheme in a single place. To achieve it, Spier and his team turned to Dynatrace.

PicPay’s partnership with Dynatrace has enabled Spier and his team to accelerate their platform engineering efforts, which has resulted in several key benefits, including the following:

  • Metrics, traces, and logs in one place. The Dynatrace platform integrates all data in one place and in context.
  • Immediate entry. Dynatrace supports most tools and languages out of the box.
    Automated notifications. The solution offers automatic alerts and the ability to create alert baselines.
  • Ease of use. All relevant data is in the same place under a single control plane with a unified view.
  • Complete visibility for Kubernetes. Since all PicPay workloads run on Kubernetes, the company can detect issues before they happen.
  • Cost efficiency. By consolidating many tools into a single platform solution, PicPay saves on costs and maintenance, preserving developer time for innovation.
To hear my full conversation with Martin Spier from PicPay, watch the customer story: Taming K8s Complexity.

The post A Kubernetes platform engineering strategy tames Kubernetes complexity appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/kubernetes-platform-engineering-tames-kubernetes-complexity/feed/ 0
Dynatrace observability now available for Red Hat OpenShift on IBM Z and LinuxONE mainframes https://www.dynatrace.com/news/blog/observability-for-red-hat-openshift-on-ibm-z-and-linuxone-mainframes/ https://www.dynatrace.com/news/blog/observability-for-red-hat-openshift-on-ibm-z-and-linuxone-mainframes/#respond Wed, 24 Jul 2024 16:59:50 +0000 https://www.dynatrace.com/news/?p=64934 OpenShift logo

Observability is fundamental when working with hybrid workloads. It serves as the foundation for informed decision-making and efficient operations. Observability provides valuable insights into system performance, resource utilization, and application health. However, achieving the end-to-end observability required to ensure the smooth operation of your systems and applications and, thus, business continuity, is challenging.

The post Dynatrace observability now available for Red Hat OpenShift on IBM Z and LinuxONE mainframes appeared first on Dynatrace news.

]]>
OpenShift logo

As we did with IBM Power, we’re delighted to share that IBM and Dynatrace have joined forces to bring the Dynatrace Operator, along with the comprehensive capabilities of the Dynatrace platform, to Red Hat OpenShift on the IBM Z and LinuxONE architecture (s390x). IBM Z and LinuxONE mainframes running the Linux operating system enable you to respond faster to business demands, protect data from core to cloud, and streamline insights and automation. By leveraging Dynatrace observability on Red Hat OpenShift running on Linux, you can accelerate modernization to hybrid cloud and increase operational efficiencies with greater visibility across the full stack from hardware through application processes.

Dynatrace full stack observability for Red Hat OpenShift

Dynatrace enhances software quality and operational efficiency, which drives innovation by unifying application, operation, and platform engineering teams on a single platform. Dynatrace with Red Hat OpenShift monitoring stands out for the following reasons:

  1. With infrastructure health monitoring and optimization, you can assess the status of your infrastructure at a glance to understand resource consumption and thus optimize resource allocation for cost efficiency.
  2. Full stack observability gives you comprehensive observability across your entire stack, including Kubernetes clusters, logs, and more. Telemetry data, such as traces and metrics, allow you to analyze the end-to-end performance of your deployed applications. You also benefit from enhanced monitoring capabilities with seamless Dynatrace integrations with cloud-native technologies and services like Istio and Prometheus.
  3. Dynatrace Smartscape® technology provides all data in context to simplify analytics and problem detection by semantically mapping metrics, traces, logs, and real user data to specific Kubernetes objects, including containers, pods, nodes, and services.
  4. You can automatically detect and analyze performance issues across your entire tech stack with Davis® AI. This lets you quickly troubleshoot issues by identifying problematic components or configurations and providing insights into the root cause.
  5. Dynatrace is designed to scale easily across the entire Kubernetes stack. Its support of GitOps practices allows for handling large-scale deployments with thousands of nodes and containers, making it the best option for enterprise-grade Red Hat OpenShift monitoring.

The new Dynatrace Kubernetes experience also marks a groundbreaking advancement for platform engineering teams by enabling you to pull critical observability and security information into a centralized Kubernetes management console. In particular, it allows you to explore cluster health, resource utilization, security, and the performance of applications built and deployed on a Kubernetes-centric platform. This is significant when coupled with the  OpenShift platform.

Clusters in Explorer in Dynatrace screenshot

Dynatrace Operator for OneAgent, API monitoring, routing, and more

Dynatrace Operator leverages Kubernetes’ native capabilities to automate the rollout, configuration, and lifecycle management of Dynatrace components within Kubernetes environments. It simplifies setting up and maintaining Dynatrace observability by encapsulating the necessary configuration and operational logic into a single entity.

Dynatrace Operator acts as an intelligent controller that understands the desired state of your Dynatrace monitoring environment and takes actions to achieve it. It automates tasks such as provisioning and scaling Dynatrace monitoring components, updating configurations, and ensuring the health and availability of your monitoring infrastructure.

Dynatrace Operator consumes DynaKubes with cloud-native full-stack configuration and deploys the following resources:

  • Dynatrace OneAgent, deployed as a DaemonSet, collects host metrics from Kubernetes nodes.
  • Dynatrace code modules, enabled via Dynatrace webhook, provide distributed tracing and code-level visibility for applications deployed on Kubernetes.
  • Dynatrace ActiveGate is used for routing and monitoring Kubernetes objects by collecting data (metrics, events, status) from the Kubernetes API.

K8s-cloud-native-architecture of Dynatrace diagram

With this approach:

  • Red Hat OpenShift infrastructure (control plane and worker nodes) and workloads are instrumented automatically without manual code change.
  • The scope of automatic workload instrumentation is user-configurable, for example, by namespace.
  • Lifecycle management of Dynatrace components (for example, ActiveGate and OneAgent) is automatic and secure.

Get started monitoring Red Hat OpenShift on IBM Z and LinuxONE

If you’re already a Dynatrace customer, refer to our Kubernetes cloud-native full stack deployment instructions.

The post Dynatrace observability now available for Red Hat OpenShift on IBM Z and LinuxONE mainframes appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/observability-for-red-hat-openshift-on-ibm-z-and-linuxone-mainframes/feed/ 0
Observe syslog with Dynatrace ActiveGate, a secure, trusted edge component https://www.dynatrace.com/news/blog/observe-syslog-with-dynatrace-activegate/ https://www.dynatrace.com/news/blog/observe-syslog-with-dynatrace-activegate/#respond Mon, 15 Jul 2024 17:18:27 +0000 https://www.dynatrace.com/news/?p=64704 Dynatrace ActiveGate

Syslog is a standard system and network device-monitoring protocol for sending event data logs to a central storage location. Integrating syslog into enterprise observability can be challenging because it doesn’t offer authentication, and syslog producers have variable flexibility and sometimes lack Transport Layer Security (TLS). The Dynatrace Environment ActiveGate edge component solves this hassle with flexible syslog endpoint configuration, making data actionable on the Dynatrace® platform.

The post Observe syslog with Dynatrace ActiveGate, a secure, trusted edge component appeared first on Dynatrace news.

]]>
Dynatrace ActiveGate

Complex syslog ecosystems can be challenging

Monitoring devices and applications that provide output via the syslog protocol is a must-have for many organizations. The key to success is making data in this complex ecosystem actionable, as many types of syslog producers exist. These include traditional on-premises network devices and servers for infrastructure applications like databases, websites, or email. You also might be required to capture syslog messages from cloud services on AWS, Azure, and Google Cloud related to resource provisioning, scaling, and security events.

Syslog’s unique nature is also often a challenge. While newer syslog daemons have implemented support for TLS encryption, you can still encounter unauthenticated and plain-text message exchanges. A local endpoint in a protected network or DMZ is required to capture these messages.

Finally, adding additional components on the edge to filter and transform syslog messages (for example, Dynatrace OpenTelemetry distribution) isn’t always possible due to architectural reasons or because it adds unnecessary complexity and cost of ownership when scaling your business.

The ultimate challenge lies in making data from syslog-supported log sources actionable. Without seeing syslog data in the context of your infrastructure, metrics, and transaction traces, you’re slowed down by manual work with siloed data. Without syslog in your observability platform, you miss out on automation, insight into service level objectives, faster mean time to repair, increased security, and resiliency.

Secure and trusted local endpoint now collects syslog data

Dynatrace now makes integrating syslog data into an AI-powered observability platform for cloud and hybrid deployments easy, simple, and secure. After ingesting syslog data safely via Environment ActiveGate, you automatically make it actionable in the right context of your infrastructure and cloud services. At the same time, you can offload the overhead and component upkeep burden to a known and trusted observability platform.

With Davis® AI automatic detection of problems and degradations in your services, you can use syslog data automatically in the context of the correct infrastructure component to fix problems faster with logs and eliminate manual correlation and guesswork. This speeds up your teams’ mean time to identify (MTTI) issues and repair (MTTR), increasing business resiliency to disruptions.

One change to send syslog to Dynatrace

You can now use the syslog ingestion endpoint on Dynatrace Environment ActiveGate for performant network and system monitoring. ActiveGate establishes Dynatrace presence in your local network and serves as a known, trusted, supported technology component with lifecycle management that greatly reduces your maintenance effort. ActiveGate also optimizes traffic volume in your network and serves as a secure relay layer in protected networks and DMZs.

Dynatrace Environmental ActiveGate syslog endpoint graphic

To enable syslog collection on an ActiveGate host, one change to extensionsuser.conf is required. No restart is necessary, and the endpoint is ready on standard ports (514 for UDP and 601 for TCP) to collect and forward logs to Dynatrace.

Change this setting in the ActivateGate extensionsuser.conf file:

#Syslog configuration
Syslogenabled=true

To complete the integration, you need to configure your syslog producer to send data to the Environment ActiveGate IP and port.

Optionally, you can adjust syslog collection beyond the default ports and configuration settings. For example, you might want to enable and provide a certificate for the TLS protocol for secure message exchange. ActiveGate uses an embedded Dynatrace OpenTelemetry Collector instance, which allows you to configure a syslog receiver according to your needs.

What’s next

See these related resources for complete details about Dynatrace syslog support:

Coming soon

  • We continue to improve syslog integration with easier failover configuration and support for monitoring configuration health.
  • Need to send syslog directly over HTTPS? More mature syslog producers support this, and we continue working on providing the required authentication layer.
Start monitoring your syslog data with the Dynatrace platform.

The post Observe syslog with Dynatrace ActiveGate, a secure, trusted edge component appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/observe-syslog-with-dynatrace-activegate/feed/ 0
Dynatrace Announces Industry’s First Observability-Driven Kubernetes Security Posture Management Solution https://www.dynatrace.com/news/press-release/dynatrace-announces-industrys-first-observability-driven-kspm-solution/ Tue, 07 May 2024 12:00:17 +0000 https://www.dynatrace.com/news/?post_type=press-release&p=63957 WALTHAM, Mass., May 7, 2024 – Dynatrace (NYSE: DT), the leader in unified observability and security, today announced it is enhancing its platform with new Kubernetes Security Posture Management (KSPM) capabilities for observability-driven security, configuration, and compliance monitoring. This announcement follows the rapid integration of Runecast technology into the Dynatrace® platform following the company’s successful […]

The post Dynatrace Announces Industry’s First Observability-Driven Kubernetes Security Posture Management Solution appeared first on Dynatrace news.

]]>
WALTHAM, Mass., May 7, 2024 – Dynatrace (NYSE: DT), the leader in unified observability and security, today announced it is enhancing its platform with new Kubernetes Security Posture Management (KSPM) capabilities for observability-driven security, configuration, and compliance monitoring. This announcement follows the rapid integration of Runecast technology into the Dynatrace® platform following the company’s successful acquisition earlier this year. The new KSPM offering builds on Dynatrace’s existing security protection capabilities, including Runtime Vulnerability Analytics (RVA) and Runtime Application Protection (RAP), strengthening cloud-native application protection in the Dynatrace platform. These combined capabilities provide DevSecOps, security, platform engineering, and SRE teams, who are all responsible for ensuring the security of Kubernetes environments, with an innovative solution for security posture and compliance.

“As workloads become more dynamic, integrating KSPM into the deployment lifecycle is essential for the security of Kubernetes environments and adherence to best practices and standards,” says KellyAnn Fitzpatrick, Senior Analyst with RedMonk. “Through a combination of real-time vulnerability assessments and contextual security insights, Dynatrace’s new KSPM solution aims to empower teams to proactively address risks and achieve complete visibility into their security posture, compliance status, and attack vectors. Such solutions can enable teams to confidently accelerate their digital transformation, knowing that their cloud-native environment is protected.”

Dynatrace’s new KSPM solution, combined with the platform’s existing RVA and RAP capabilities, enables teams to instantly detect risks in their Kubernetes deployments and automatically prioritize remediation based on risk exposure. To achieve this at scale, Dynatrace uses Davis® hypermodal AI, which combines predictive and causal AI techniques for precise answers and automation and generative AI for increased productivity. As a result, teams can gain comprehensive insights into cloud-native applications across code, libraries, language runtime, and container infrastructure. This helps them strengthen the security posture of their Kubernetes deployments, while also complying with regulatory frameworks and industry best practices, including those defined by the Center for Internet Security (CIS).

“While most teams rely on agentless workload scanning to enable KSPM, this snapshot approach is incapable of providing real-time insights and often causes a frequency of excessive alerts, without runtime context, that waste time and make remediation efforts more confusing,” said Bernd Greifeneder, CTO at Dynatrace. “As the attack surface of cloud applications continues to increase and attackers are targeting the misconfiguration of cloud infrastructure, APIs, and the software supply chain itself, organizations must evolve their solutions and processes to ensure their cloud technologies, including Kubernetes, are constantly protected against the latest risks and empower teams to operate with best practices. Building on the Dynatrace platform’s existing RVA and RAP capabilities, Dynatrace is delivering an innovative approach to KSPM that we believe sets us apart, providing teams with real-time insights, automated compliance, and actionable advice that free up resources and reduce exposure risk.”

Dynatrace Kubernetes Security Posture Management is expected to be generally available in the second half of 2024.

Cautionary Language Concerning Forward-Looking Statements

This press release includes certain “forward-looking statements” within the meaning of the Private Securities Litigation Reform Act of 1995, including statements regarding the capabilities of Dynatrace Kubernetes Security Posture Management, and the expected benefits to organizations from using Dynatrace Kubernetes Security Posture Management. These forward-looking statements include all statements that are not historical facts and statements identified by words such as “will,” “expects,” “anticipates,” “intends,” “plans,” “believes,” “seeks,” “estimates,” and words of similar meaning. These forward-looking statements reflect our current views about our plans, intentions, expectations, strategies, and prospects, which are based on the information currently available to us and on assumptions we have made. Although we believe that our plans, intentions, expectations, strategies, and prospects as reflected in or suggested by those forward-looking statements are reasonable, we can give no assurance that the plans, intentions, expectations, or strategies will be attained or achieved. Actual results may differ materially from those described in the forward-looking statements and will be affected by a variety of risks and factors that are beyond our control, including the risks set forth under the caption “Risk Factors” in our Quarterly Report on Form 10-Q filed on February 8, 2024, and our other SEC filings. We assume no obligation to update any forward-looking statements contained in this document as a result of new information, future events, or otherwise.

The post Dynatrace Announces Industry’s First Observability-Driven Kubernetes Security Posture Management Solution appeared first on Dynatrace news.

]]>
Stream logs to Dynatrace with Amazon Data Firehose to boost your cloud-native journey https://www.dynatrace.com/news/blog/stream-logs-to-dynatrace-with-amazon-data-firehose/ https://www.dynatrace.com/news/blog/stream-logs-to-dynatrace-with-amazon-data-firehose/#respond Fri, 03 May 2024 15:25:51 +0000 https://www.dynatrace.com/news/?p=63910 Dynatrace and Amazon Data Firehose

Your cloud logs can provide the root cause of high-impact issues or reveal the details of security incidents. Now, you can integrate an Amazon Data Firehose high-frequency data stream directly with the high-performant Dynatrace Grail™ analytics engine and use the Dynatrace AI-powered observability platform to mitigate issues with minimal impact to your business.

The post Stream logs to Dynatrace with Amazon Data Firehose to boost your cloud-native journey appeared first on Dynatrace news.

]]>
Dynatrace and Amazon Data Firehose

Real-time streaming needs real-time analytics

As enterprises move their workloads to cloud service providers like Amazon Web Services, the complexity of observing their workloads increases. Log data—the most verbose form of observability data, complementing other standardized signals like metrics and traces—is especially critical. As cloud complexity grows, it brings more volume, velocity, and variety of log data.

Managing this change is difficult. Without the ability to see the logs that are relevant to your service, infrastructure, or cloud function—at exactly the right time and in exactly the right format—your cloud or DevOps engineers lose the ability to find the root causes of the issues they troubleshoot. Even the AIOps approach doesn’t cut it if you don’t have proper logs in your observability platform.

Amazon CloudWatch is the most common method of collecting logs across your AWS footprint. As a native tool used by many enterprises, CloudWatch supports a wide range of AWS resources, applications, and services.

Amazon Data Firehose helps stream logs to the right destination

But your SREs and DevOps engineers know CloudWatch is not the terminal destination for data but rather an intermediate station. Their job is to find out the root cause of any SLO violations, ensure visibility into the application landscape to fix problems efficiently and minimize production costs by reducing errors. SREs and DevOps engineers need cloud logs in an integrated observability platform to monitor the whole software development lifecycle.

When trying to address this challenge, your cloud architects will likely choose Amazon Data Firehose. This fully managed native service is indispensable for streaming high-frequency logs collected by CloudWatch.

In some deployment scenarios, you might skip CloudWatch altogether. Take the example of Amazon Virtual Private Cloud (VPC) flow logs, which provide insights into the IP traffic of your network interfaces. VPC flow logs can be used as the source for troubleshooting connectivity issues, implementing security incident investigations, detecting intrusions, or managing access control issues. VPC flow logs can be massive in volume as your cloud deployment footprint grows, and directly streaming these logs with Amazon Data Firehose can be the most cost-effective method.

After configuring Amazon Data Firehose, your teams discover they have completed only the first part of the observability jigsaw puzzle. They also need a high-performance, real-time analytics platform to make that data actionable.

Dynatrace delivers the missing piece for AWS cloud observability with native Firehose integration. This complements our existing AWS logging integrations like S3 log forwarder, Lambda layer log forwarding, or direct log ingest API. These already provide a common integration with AWS log sources. The new Firehose integration removes intermediary components that previously required additional maintenance and provides a direct link from AWS to Grail data lakehouse.

This means high-frequency streamed logs from Firehose can be captured in your Dynatrace environment, automatically processed, stored in Grail for the retention period of your choice, and included in the full observability automation suite of the Dynatrace® platform, apps, and Davis® AI problem detection.

With this out-of-the-box support for scalable data ingest, log data is immediately available to your teams for troubleshooting and observability, investigating security issues, or auditing. As logs are first-class citizens alongside traces, metrics, business events, and other data types, you have an observability platform ready to scale with you in your cloud-native journey.

Easy setup takes just a few steps

Setting up a direct ingest of Firehose log data is quick and easy.

First, you need to generate an API key to ingest logs. In the Dynatrace web UI, go to Access tokens and select Generate new token. Select ingest logs as the scope of the token. Then, generate the token.

Next, go to the AWS console to configure the forwarding of data streams defined in your log groups. Data Firehose stream requires a trusted relationship with CloudWatch through an IAM role. Follow the instructions available in Dynatrace documentation to allow proper access and configure Firehose settings.

Now, you can set up your Firehose stream. The preferred way is to use a CloudFormation template that streamlines and automates the process. See CloudFormation template documentation for details.

Alternatively, you can configure the stream in the AWS web console. Choose Dynatrace as the Destination in the AWS console and complete the other fields with the correct parameters.

Choose Dynatrace as the destination in AWS console.
Figure 1. Choose Dynatrace as the destination in AWS console.

Now, you can view your cloud logs in Dynatrace!

For example, open the Clouds app with integrated logs in the context of your Lambda functions observability for one-click access to error logs.

See logs in context in the Dynatrace Platform, including the relevant logs for this AWS Lambda function.
Figure 2. See logs in context in the Dynatrace Platform, including the relevant logs for this AWS Lambda function.

When Dynatrace Davis AI detects a problem in your environment, you can also see relevant logs streamed via AWS Firehose that are related to the problem. When analyzing a problem, look at the related service, which displays related log data. This lets you jump right to the error that provides details of the problem.

When doing proactive health checks or analysis, you can inspect log data in Notebooks. For example, pick a template to explore data or write your own DQL query and chart incoming error rates from logs streamed via AWS Firehose.

Easily visualize Lambda error log distribution over time with Notebooks.
Figure 3. Easily visualize Lambda error log distribution over time with Notebooks.

Try it out today

Share your experience

We’d love to hear from you. Share your use cases for Amazon Data Firehose integration with the Dynatrace Community.

Stream AWS service logs collected in CloudWatch or directly via Firehose.

The post Stream logs to Dynatrace with Amazon Data Firehose to boost your cloud-native journey appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/stream-logs-to-dynatrace-with-amazon-data-firehose/feed/ 0
Enable full observability for Linux on IBM Z mainframe now with logs https://www.dynatrace.com/news/blog/enable-full-observability-for-linux-on-ibm-z-mainframe-now-with-logs/ https://www.dynatrace.com/news/blog/enable-full-observability-for-linux-on-ibm-z-mainframe-now-with-logs/#respond Wed, 17 Apr 2024 17:25:45 +0000 https://www.dynatrace.com/news/?p=63665 Hosts fetch logs

Complement your hybrid cloud journey with a resilient observability setup with log monitoring on Linux on IBM Z and LinuxONE. Include the familiar mainframe OS into one integrated observability platform and thus eliminate the need for platform-specific component upkeep and management, reduce the risk of prolonged outages, and get full details from logs for troubleshooting.

The post Enable full observability for Linux on IBM Z mainframe now with logs appeared first on Dynatrace news.

]]>
Hosts fetch logs

Mainframe is a strong choice for hybrid cloud, but it brings observability challenges

IBM Z is a mainframe computing platform chosen by many organizations with a hybrid cloud strategy because of its security, resiliency, performance, scalability, and sustainability. With the availability of Linux on IBM Z and LinuxONE, the IBM Z platform brings a familiar host operating system and sustainability that could yield up to 75% energy reduction compared to x86 servers.

That’s why a hybrid cloud scenario, where workloads are shared between public clouds and highly performant mainframe platforms like IBM Z, is a robust and effective strategy.

The challenge for hybrid cloud deployments is maintaining critical observability, which must include the full set of monitoring signals: logs, metrics, and traces. Without combining these signals in a unified AI-powered observability platform, monitoring apps, infrastructure, and troubleshooting issues are nothing more than a patchwork of manual correlation.

Deploying your critical applications on additional host operating systems increases the dependencies for observability. It means maintaining platform-specific observability components or tools, managing security updates, and deploying changes, all of which lead to configuration spread.

This creates a risk that can impact your time to problem resolution in troubleshooting, the effectiveness of AIOps workflows to remediate issues before they affect your end-users, and ultimately your business metrics.

Logs become an integrated part of observability

Dynatrace provides a unified and integrated platform to observe such hybrid cloud deployments, now with added support to monitor logs effortlessly on Linux on IBM Z and LinuxONE.

OneAgent® is a core component of the Dynatrace platform; it enables observability with minimal setup effort while offering extensive and flexible central configuration options. By including logs in your hybrid cloud observability, you have everything you need in one place to make smarter, faster decisions when troubleshooting and measuring the health of your application environments.

You can now seamlessly expand your analysis of the root cause of any problem identified by Davis® AI with logs automatically available in the correct context of hosts, applications, or other identifiers specific to your environment.

Because Dynatrace provides a unified and central place to configure your observability, there is a single place where you manage your log collection for public and private clouds and mainframe components like Linux on IBM Z or LinuxONE.

This makes your log collection policies much more effective and transparent. You can push a filtering change to filter out all unwanted logs from your central Dynatrace environment and apply the change automatically to all your monitored platforms.

It’s also easy to minimize the risk of violating data access policies or regulations by masking sensitive data in logs. By centrally configuring masking rules for sensitive data, you can push the rules out to all of your deployed OneAgents wherever they are deployed to make sure you stay compliant.

Configure log collection across all your hosts

Start by deploying OneAgent for Linux on IBM Z by going to Deploy OneAgent (for earlier versions of Dynatrace and Dynatrace Managed deployments, go to Settings > Deploy Dynatrace), choose Linux as the underlying platform, and s390 as the installer type.

You can now install OneAgent on Linux with s390 architecture.
Figure 1. You can now install OneAgent on Linux with s390 architecture.

Next, set up log ingest. As log monitoring is now available with OneAgent for Linux on IBM Z, a single log ingest rule can cover all your Linux operating systems no matter what architecture is utilized under the hood. This means OneAgents deployed on Linux with s390, ARM, AIX, or x86 are covered.

Go to Settings > Log Monitoring > Log ingest rules and turn on Ingest all logs to start log collection.

Enabling the log ingestion setting applies to deployed OneAgents on Linux on IBM Z and LinuxONE.
Figure 2. Enabling the log ingestion setting applies to deployed OneAgents on Linux on IBM Z and LinuxONE.

Next you can start using logs in your troubleshooting and analysis tasks. For example, on the Dynatrace platform, open the new Infrastructure & Operations app and navigate to any monitored host running on Linux on IBM Z (s390 architecture). You can see the Logs tab for the host, which displays insights about automatically contextualized logs from that host.

Infrastructure & Operations app shows a monitored host with s390 architecture, and the Logs tab shows log data for that host.
Figure 3. The infrastructure & Operations app shows a monitored host with s390 architecture, and the Logs tab shows log data for that host.

You can take your Dynatrace Grail™ analysis of log data further in Notebooks. Start with a query builder to get error logs for your Linux on IBM Z hosts, and continue your exploration with Dynatrace Query Language.

Error logs in Notebooks with distribution chart
Figure 4. Error logs in Notebooks with distribution chart

Start monitoring logs on Linux on IBM Z

What’s next

Stay tuned for an upcoming blog post about log collection in OpenShift for Linux on IBM Z and LinuxONE.

Are you running containerized applications on IBM Z?

The post Enable full observability for Linux on IBM Z mainframe now with logs appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/enable-full-observability-for-linux-on-ibm-z-mainframe-now-with-logs/feed/ 0
The 3 biggest Kubernetes deployment mistakes you can make https://www.dynatrace.com/news/blog/the-3-biggest-kubernetes-deployment-mistakes/ https://www.dynatrace.com/news/blog/the-3-biggest-kubernetes-deployment-mistakes/#respond Tue, 02 Apr 2024 14:55:42 +0000 https://www.dynatrace.com/news/?p=63292 Kubernetes deployment graphic

The last decade brought a wave of digital transformation that accelerated the move towards cloud-native tech, specifically Kubernetes. We could talk about this move’s benefits (and downsides!) for more than a few blogs, but that’s not why we’re here. We’re here to talk about the less savory side of cloud-native transformation—you know, all the things […]

The post The 3 biggest Kubernetes deployment mistakes you can make appeared first on Dynatrace news.

]]>
Kubernetes deployment graphic

The last decade brought a wave of digital transformation that accelerated the move towards cloud-native tech, specifically Kubernetes. We could talk about this move’s benefits (and downsides!) for more than a few blogs, but that’s not why we’re here. We’re here to talk about the less savory side of cloud-native transformation—you know, all the things that can go wrong. More specifically, all the Kubernetes deployment mistakes you can make when moving to Kubernetes—and how to avoid making them in the first place.

As someone who has worked deep in the coding trenches with developers my whole life, I’ve hand-picked the top three mistakes you can make when moving to Kubernetes. So, without further ado, let me share these hard-earned mistakes you should avoid like the plague when moving to Kubernetes!

Kubernetes deployment mistake #1: Managing Kubernetes from the command line

Kubernetes deployments almost feel like magic the first time you get them working. You use a (hopefully) short YAML file to specify the application you want to run, and Kubernetes just makes it so. Make a change to the file, apply it, and it will update in near real-time.

But as powerful as kubectl is, and as instructive as it can be to explore Kubernetes using it, you should not come to rely on kubectl too much. Of course, you’ll return to it (or its amazing cousin, k9s) when you need to troubleshoot issues in Kubernetes, but don’t use it to manage your cluster.

Kubernetes was made for the configuration-as-code paradigm, and all those YAML files belong in a Git repo. You should commit any and all of your desired changes to a repo and have an automated pipeline deploy the changes to production. Some of your options include:

Kubernetes deployment mistake #2: Forgetting all about resources

Let’s assume all your workloads are up and running with all the goodness of Kubernetes and configuration as code. But now you’re orchestrating containers, not virtual machines. How do you ensure they get the CPU and RAM they need? Through resource allocation!

Resource requests

What happens if you forget to set resource requests?

Kubernetes will pack all your pods (“workloads” in Kubernetes-speak) into a handful of nodes. They won’t get the resources they need. The cluster won’t scale itself up as needed.

What are resource requests?

Resource requests tell the scheduler how many resources you expect your application to consume. When assigning pods to nodes, Kubernetes budgets them so that the node’s resources meet all of their requirements.

Resource limits

What happens if you forget to set resource limits?

A single pod may consume all the CPU or memory available on the node, causing its neighbors to be starved of CPU or hit Out of Memory errors.

What are resource limits?

Resource limits let the container runtime know how many resources you allow your application to consume. For the CPU limit, your application will be able to get that much CPU time but no more. Unfortunately (for the application), if it hits the memory limit, it will be OOMKilled by the container runtime.

So, go ahead and define requests and limits for each of your containers. If you aren’t sure, just take a guess, and keep in mind that the safe side is higher. Whether you’re certain or not, make sure to monitor actual resource usage by your pods and containers by using your cloud provider or APM tools.

Kubernetes deployment mistake #3: Leaving the developers behind

Immutable infrastructure and clean upgrades. Easy scalability. Highly available, self-healing services. Kubernetes provides you with lots of value directly out of the box. Unfortunately, this value might not be a priority for the developers working on your product. Your developers have other concerns:

  • How do I build and run my code?
  • How do I understand what my code is doing in development, testing, and integration?
  • How do I investigate bugs reported in QA and production environments?

For many of these tasks, Kubernetes pulls the rug out from under the developer. Running development environments locally is much harder because many dev and test workloads are moved to the cloud. The code-level visibility developers rely on is often poor in these environments, and direct access to the application and its filesystem is virtually impossible.

Successful Kubernetes adoption requires the right tools

To lead a successful adoption of a new platform such as Kubernetes, you need everyone to see the value in it. But don’t forget that developers require the right tools to keep up with their code and understand what it’s doing as it’s running.

Get started on your Kubernetes journey with Dynatrace.

The post The 3 biggest Kubernetes deployment mistakes you can make appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/the-3-biggest-kubernetes-deployment-mistakes/feed/ 0
Bring syslog into Dynatrace using OpenTelemetry to get open source value with enterprise support https://www.dynatrace.com/news/blog/bring-syslog-into-dynatrace-over-opentelemetry-to-get-open-source-value-with-enterprise-support/ https://www.dynatrace.com/news/blog/bring-syslog-into-dynatrace-over-opentelemetry-to-get-open-source-value-with-enterprise-support/#respond Fri, 15 Mar 2024 16:45:45 +0000 https://www.dynatrace.com/news/?p=63075 Fetch logs

Syslog is a standard protocol for system and network device monitoring. Integrating syslog into enterprise observability solutions is tricky due to its strict support and security patching requirements.

The new Dynatrace OTel Collector distribution unlocks the power of syslog and open source community contributions with the power of Dynatrace support and the value of Dynatrace Grail™ to analyze log data from devices at scale.

The post Bring syslog into Dynatrace using OpenTelemetry to get open source value with enterprise support appeared first on Dynatrace news.

]]>
Fetch logs

Getting insights into the health and disruptions of your networking or infrastructure is fundamental to enterprise observability. Syslog is the go-to protocol that delivers infrastructure administrators, network engineers, and security team logs that tell them all they need to know about their systems’ delivery, performance, availability, and security.

Without syslog, you’re blind to what happens on your infrastructure

While syslog is a common way to gain insights into enterprise infrastructure operations, integrating it with other signals into an observability overview is often a painful experience.

Syslog is a protocol with clear specifications that require a dedicated syslog server. This is needed to collect messages across your systems because many different types of devices and applications can produce logs in the syslog format.

However, enterprise adoption at scale typically has much higher requirements for components than for supported features—components must have proper vendor support. Without vendor support, you’re betting your business on goodwill. Even for a supported component, delivering logs from applications and infrastructure to DevSecBizOps workflows requires significant manual configuration.

For example, a supported syslog component must support the masking of sensitive data at capture to avoid transmitting personally identifiable information or other confidential data over the network. Log batching, enrichment, transformation, log source distinction, and application offloading are also regular requirements.

As enterprise environments scale enormously, filtering and dropping data “at the edge” before transmission to a central collection point must be a supported option.

Compliance, retention, archiving, or data governance regulations often require multicasting logs from the original source to multiple destinations, like an observability platform with long-term log storage.

In the end, site reliability engineering (SRE) and security teams need to have data delivered via syslog to their observability platform, in the context of other data types.

Syslog will remain a proven log solution because without understanding why connections are dropped, server starts or stops, or which requests your firewall blocked, your organization runs like a ship where the captain on the bridge has no understanding of what’s going on at the lower levels of the ship. However the challenges in maintaining syslog in a cloud-native era create a maze of requirements that SRE teams and infrastructure administrators must navigate, often finding themselves maintaining multiple tools and components.

This increases the risk of multiple points of failure, adds overhead, and ultimately fractures observability overview with prolonged time needed to recover from potential outages.

Start monitoring syslog using OpenTelemetry under the Dynatrace umbrella of support

OpenTelemetry has been a rising star in the observability landscape and is often a preferred way to achieve end-to-end visibility with a vendor-agnostic footprint. Dynatrace has been a part of the OpenTelemetry journey for years and has contributed to its rise.

With the new Dynatrace OTel Collector distribution, we provide a streamlined and supported way to collect logs using the syslog protocol. This fills all the requirements enterprises have and makes it hassle-free to stream syslog to Grail data lakehouse integrating logs with other observability data.

The Dynatrace OTel Collector for syslog has numerous benefits. Our approach is to understand what components our customers need and value. We then integrate them with our observability platform and offer support, so you don’t have to worry about unsupported bugs or lack of ownership.

We also provide security updates and patches to critical vulnerabilities that may arise in the components. This alleviates the risk of open source components with unpatched vulnerabilities remaining open to exploitation long after they have been revealed.

Ultimately this combination of Dynatrace support and the OpenTelemetry standard gives you the best of both worlds—enterprise-grade software support with open source community contributions.

Dynatrace OTel Collector fits with your existing setup

The new Dynatrace OTel Collector fits nicely into your existing Dynatrace setup to bring in syslog data. Our existing log ingest API already supports your logs using the OpenTelemetry protocol, so you just need to deploy the collector and point your syslog producers to it.

To start using the Dynatrace OTel Collector, take the following steps:

  1. Generate an API token for the OTLP endpoint in your environment.
  2. Find our newly released Dynatrace OTel Collector, deploy it, and configure the exporter with your API key and environment ID.
  3. Configure receivers to enable different log sources for your syslog producers.
  4. Point your syslog sources to the collector and you’re done!

This diagram explains how the components communicate with each other.

This diagram explains how the components of the Dynatrace OTel Collector communicate with each other.

Take a look at this example for configuration. After generating an API token and deploying the collector, configure your instance. You need to configure each component (receiver, optional processor, and exporter) individually in a YAML file and enable them via pipelines. Follow the examples below or refer to Collector configuration documentation.

To point the exporter to your environment’s OTLP endpoint, add the following configuration:

exporters:
  logging:
    verbosity: detailed

  otlphttp/tenant_1:
    endpoint: "https://{your-tenant}.live.dynatrace.com/api/v2/otlp"
    headers:
      Authorization: "Api-Token {your-api-token}"

Next, you can add receivers to your collectors, for example, F5 BIG-IP systems to log to a remote syslog server (version 11.x-17.x). Refer to F5 BIG-IP documentation for detailed and up-to-date instructions regarding remote Syslog configuration. Take a look at Syslog (Dynatrace OTel Collector) in Dynatrace Hub for an example configuration file for the receiver, so you can enable two separate syslog endpoints for F5 and host syslogs. This allows you to differentiate log sources (attribute.device.type) for analysis in Dynatrace.

You can also make the Dynatrace OTel Collector multicast incoming syslog messages to multiple destinations. For example, you can set up exporters for your Dynatrace production environment and sandbox environment:

service:
  pipelines:
    logs:
      receivers: [syslog/f5, syslog/host]
      processors: [batch]
      exporters: [logging, otlphttp/tenant_1, otlphttp/tenant_2]

As a result, you should see logs in Dynatrace with corresponding log.source and device.type attributes:

Logs in Dynatrace with corresponding log.source and device.type attributes

Deploy Dynatrace OTel Collector for syslog now

What’s next

  • Stay tuned for direct syslog ingestion into Dynatrace, which brings syslog endpoints to an Environment ActiveGate, fully configurable from the cluster. This will enable you to use Dynatrace ActiveGate to ingest syslog data.

Go to Syslog (Dynatrace OTel Collector) in Dynatrace Hub to see examples and continue to the installation.

The post Bring syslog into Dynatrace using OpenTelemetry to get open source value with enterprise support appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/bring-syslog-into-dynatrace-over-opentelemetry-to-get-open-source-value-with-enterprise-support/feed/ 0
Dynatrace strengthens container security across popular cloud-based registries https://www.dynatrace.com/news/blog/dynatrace-strengthens-container-security-across-popular-cloud-based-registries/ https://www.dynatrace.com/news/blog/dynatrace-strengthens-container-security-across-popular-cloud-based-registries/#respond Wed, 06 Mar 2024 21:27:30 +0000 https://www.dynatrace.com/news/?p=62925 Container security graphic

Cosign-signed, immutable images for cloud and Kubernetes environments

The post Dynatrace strengthens container security across popular cloud-based registries appeared first on Dynatrace news.

]]>
Container security graphic

Cloud-native CI/CD pipelines and build processes often expose Kubernetes to attack vectors via internet-sourced container images. Despite scanning, these container images can still be susceptible to supply chain attacks if not properly verified. Ensuring immutability—coupled with thorough scanning and strict verification—is crucial for any container entering Kubernetes clusters. This is particularly vital for securing observability solutions like Dynatrace® Kubernetes infrastructure observability, application observability, and Application Security.

The Dynatrace Operator is responsible for the secure lifecycle of components necessary for Kubernetes cluster monitoring. Dynatrace Operator ensures secure download and rollout of components via protected connections to the Dynatrace platform. Introducing Cosign-signed immutable images, Dynatrace further empowers you to independently verify images, maintaining observability free from supply chain attacks.

The benefits of independently verifiable container images begin, but do not end, with enhanced security.

  • Security: Signing and immutability of container images significantly reduce the risk of security breaches, ensuring that only verified, tamper-proof observability tools are deployed.
  • Compliance: Adhering to stringent security standards helps meet regulatory and compliance requirements for cloud-native environments.
  • Reliability: Immutable images guarantee consistent performance and behavior, enhancing stability.

Incorporating signed Dynatrace containers into your pipeline

To enhance security in CI/CD processes, Dynatrace customers can integrate verified Dynatrace container images into their deployment pipelines. This process, illustrated below using Amazon’s Elastic Container Registry (ECR), is also applicable to Docker Hub and will be extended to other cloud platforms in the future.

How it works

Begin by browsing Dynatrace images on the Amazon Elastic Container Registry (ECR) Public Container Gallery. Signed and immutable container images are available for the entire Dynatrace observability stack.

  • Dynatrace OneAgent®, which offers Application Security, host, network, and infrastructure monitoring
  • Dynatrace OneAgent code modules, which offer application observability and security for Go, Java, Node, .Net, and PHP
  • Dynatrace Operator
  • Dynatrace ActiveGate for Kubernetes monitoring, routing, and signal compression

Drilling down on a container image reveals tags for both signatures as well as versions of the binaries. The version numbers of these binaries correlate 1:1 with the container image’s tags. This correlation ensures that Dynatrace software components are versioned exactly the same way for both containerized and non-containerized workloads.

Using public images

The recommended way of using Dynatrace signed immutable images is to use a private registry, which is typical for most organizations looking to prioritize security in their CI/CD workflows. This approach offers potentially improved performance and reliability, as the registry can be optimized for specific network environments.

This process involves a few steps:

  1. Query your cluster on latest available OneAgent, code module, and ActiveGate tag information.
  2. Copy container image to private registry from public registry.
  3. Check that the images are valid and secure.
  4. Reference the container image in the DynaKube.

To determine the latest supported version of the Dynatrace OneAgent container image execute the following command to get the JSON output from the REST API:

curl https://{your_tenant}.live.dynatrace.com/api/v1/deployment/image/agent/oneAgent/latest -
-header 'Authorization: Api-Token {your_API_token}'

Required token scope is the following: InstallerDownload

Here is the example of the the reply to the provided REST API call is in the JSON format:

{
"source": "{public_ecr_address}/dynatrace/dynatrace-oneagent",
"tag": "1.279.242.20240108-114943"
}

The retrieved image tag can now be used to copy the container image from the provided ECR to your private image registry.

To copy all image architectures and Cosign signatures for verification, make sure to use the --all flag and set use-sigstore-attachments to true in Skopeo’s container registry configuration.

skopeo copy --all docker://public.ecr.aws/dynatrace/dynatrace-oneagent:<tag> \
docker://registry.my-company.com/dynatrace-oneagent:<tag>

Finally, verify the container image signature to ensure authenticity and integrity.

cosign verify --insecure-ignore-tlog --key https://ca.dynatrace.com/v1/cosign.pub \

  registry.my-company.com/dynatrace-oneagent:<tag>

Once the image has been copied over to your private registry, it can be scanned for vulnerabilities.

Finally, reference the verified private registry images in the DynaKube. The example below shows all three images: OneAgent, CodeModules, and ActiveGate. Note the inclusion of a pull secret, required for protected private registries.

apiVersion: dynatrace.com/v1beta1

kind: DynaKube

metadata:

  name: private-registry

  namespace: dynatrace

spec:

  apiUrl: https://(your environment)/api

  tokens: api-tokens

  customPullSecret: pull-secret

  oneAgent:

    cloudNativeFullStack:

      image: (your-registry)/dynatrace-oneagent:1.279.242.20240108-114943

      codeModulesImage: (your-registry)/dynatrace-codemodules:1.279.242.20240108-114943

  activeGate:

    capabilities:

      - kubernetes-monitoring

      - routing

    image: (your-registry)/dynatrace-activegate:1.279.116.20231206-155926

Take cloud-native security to new heights

The latest release of Dynatrace signed immutable container images marks a step forward in securing cloud-native observability stacks. By ensuring the integrity and security of containers at every step, from public registry to Kubernetes deployment, Dynatrace sets a new standard in cloud-native security. This gets to the heart of the Dynatrace mission to provide unparalleled observability and security in an ever-evolving digital landscape.

Video thumbnail

What’s next

New users can explore these advanced features with a free 15-day trial, experiencing firsthand how Dynatrace is transforming cloud-native security.

For current Dynatrace customers, getting started with our new signed, immutable images is easy—just refer to Dynatrace Documentation.

The post Dynatrace strengthens container security across popular cloud-based registries appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-strengthens-container-security-across-popular-cloud-based-registries/feed/ 0
How platform engineering and IDP observability can accelerate developer velocity https://www.dynatrace.com/news/blog/how-platform-engineering-can-accelerate-developer-velocity/ https://www.dynatrace.com/news/blog/how-platform-engineering-can-accelerate-developer-velocity/#respond Wed, 06 Mar 2024 20:40:49 +0000 https://www.dynatrace.com/news/?p=62869 Data privacy by design, CrowdStrike

At Dynatrace Perform 2024, Dynatrace colleagues Andreas Grabner and Adam Gardner discussed how platform engineering accelerates developer velocity.

The post How platform engineering and IDP observability can accelerate developer velocity appeared first on Dynatrace news.

]]>
Data privacy by design, CrowdStrike

As organizations look to expand DevOps maturity, improve operational efficiency, and increase developer velocity, they are embracing platform engineering as a key driver. Indeed, recent research found that 54% of organizations are investing in platforms to enable easier integration of tools and collaboration between teams involved in automation projects.

Platform engineering creates and manages a shared infrastructure and set of tools, such as internal developer platforms (IDPs), to enable software developers to build, deploy, and operate applications more efficiently. The goal is to abstract away the underlying infrastructure’s complexities while providing a streamlined and standardized environment for development teams. As a result, teams can focus on writing code and building features, rather than dealing with infrastructure nuances.

During a breakout session at Dynatrace Perform 2024, Dynatrace DevSecOps activist Andreas Grabner and staff engineer Adam Gardner demonstrated how to use observability to monitor an IDP for key performance indicators (KPIs). The pair showed how to track factors, including developer velocity, platform adoption, DevOps research and assessment metrics, security, and operational costs.

Recent Dynatrace research has found that only 40% of a typical engineer’s time is spent on productive tasks, and 36% of developers resign because of a bad developer experience, Grabner noted. “If your developers are leaving the company, the IDP may have something to do with it,” he said.

Platform engineering: Build for self-service

Self-service deployment is a key attribute of platform engineering. It gives developers the means to create environments and toolsets unique to their projects.

“[An IDP] must be a product that developers want to use because it helps them get the job done,” Grabner said. “It makes them more productive . . . and reduces the complexity of things such as reading a new app or service. They shouldn’t worry about the platform; they should just start writing code.”

Because of their versatility, teams can use IDPs for all types of software engineering projects, not just those in cloud-native scenarios. IDPs can eliminate much of the administrative minutiae that stalls development projects. Grabner gave the example of one Dynatrace banking customer who built an IDP that enables developers to provision new Microsoft Azure machines or Chef policies without administrative help. “IDPs are not constrained to building microservices or a new serverless app,” Grabner noted.

Before putting an IDP in place, organizations must encourage their platform engineering teams to adopt a product mindset with feedback loops between developers and users. They should also establish milestones to ensure the built product solves a defined business problem.

Reference IDP with Dynatrace

The Dynatrace IDP encompasses platform services, delivery services, and access to observability and automation tools. The Dynatrace Operator automatically ingests all observability data from OpenTelemetry and Prometheus. Furthermore, OneAgent® software observes and gathers all remaining workload logs, metrics, traces, and events.

Automate deployment for faster developer velocity

Additionally, the IDP used during the session connects to the open source Backstage developer portal platform and a library of templates stored in a GitLab repository. The templates can deploy automatically into the development environment with just a few clicks.

Argo works in a GitOps fashion to automate the deployment of files stored in Git. “Argo has an eagle eye on the Git repository,” Gardner said. “Every time something changes, it’s synced to Kubernetes.”

Backstage holds many of an organization’s critical development resources that must be treated with the same respect as business-critical data. Observability is not only about measuring performance and speed but also about capturing granular business analytics to support data-driven decision-making. These metrics can include how many people are using the IDP, how quickly the tasks are running in the IDP, and more. “That means making it available, resilient, and secure,” Grabner said.

Intelligent monitoring is also crucial. “If you don’t monitor, you risk building a product that nobody needs,” Grabner continued.

Observability is a critical component of an IDP. It illuminates the activity of components such as Backstage, GitHub, Argo, and other tools. Service-level objectives (SLOs) are similarly important. SLOs help developers to accelerate their velocity and remain productive with an optimally functioning platform.

Test continuously

Synthetic testing simulates user behaviors within an application or service to pinpoint potential problems. This process is vital to an IDP’s effectiveness. An observability solution can monitor both synthetic and real-user tests to verify an application is on track.

GitLab, a source code repository and collaborative software development platform for DevOps and DevSecOps projects, is populated with a set of pre-filled templates. The combination gives developers a unique set of tools they can deploy on a self-service basis with full monitoring by Dynatrace.

“Every time [developers] pick a template in Backstage, they get their own version of the Git repository based on the template. Then, Argo deploys the app,” Grabner said. “It has worked kind of flawlessly.”

Observability at the core

How we built the IDP

Platform engineering is about being responsible for making sure platforms are available,” Gardner said. “Dynatrace can tell us whether Argo is up and whether it’s killing GitHub with too many syncs. It lets us see events such as starts and traces in a standardized manner.” This certainty can accelerate developer velocity and improve the developer experience, resulting in better software and happier, more productive developers.

Dynatrace has made the reference IDP architecture available on GitHub for anyone to use. It includes a notebook with configuration and deployment instructions.

“It explains every single step that was involved in building the IDP, creating the configuration, and setting up Argo,” Gardner said. “You can launch a code space that starts a container that shows you everything about how an app was built and deployed.”

Curious to learn more about observability to optimize KPI success? Check out the Perform 2024 session: Observability guide to platform engineering.

FAQs about platform engineering

What is the primary role of an internal developer platform (IDP)?

An IDP is a shared infrastructure and set of tools created and managed by platform engineering teams. By providing a self-service, standardized environment, IDPs enable software developers to more efficiently build, deploy, and operate applications. IDPs reduce the need for developers to manage underlying infrastructure nuances, thereby simplifying workflows and accelerating application delivery.

How does platform engineering improve software development?

Platform engineering accelerates developer velocity by providing a streamlined, standardized environment that abstracts away infrastructure complexities. This allows developers to focus on writing code and building features more efficiently.

It improves the developer experience by offering self-service tools and automating tasks, reducing the cognitive load and administrative minutiae that can otherwise hinder productivity and frustrate developers.

Why is observability crucial for successful platform engineering?

Observability is crucial for successful platform engineering because it provides deep insights into the performance, health, and activity of the internal developer platform and its components. It allows platform teams to monitor key performance indicators (KPIs) like platform adoption and operational costs, helping to support the platform’s availability, resiliency, and security. Continuous monitoring helps verify the effectiveness of the IDP and ensures developers maintain optimal productivity.

What are the key principles for building an effective internal developer platform?

Building an effective internal developer platform involves adopting a product mindset, treating developers as internal customers, and incorporating their feedback through continuous loops. Here are some principles to keep in mind as you build:

  • Provide clear self-service capabilities.
  • Automate repetitive tasks.
  • Focus on solving common developer pain points.

The platform should also emphasize scalability, security, and compliance, while fostering a culture of collaboration and knowledge sharing within the engineering organization.

Discover how unified observability unlocks platform engineering success in the free ebook: Driving DevOps and platform engineering for digital transformation.

The post How platform engineering and IDP observability can accelerate developer velocity appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/how-platform-engineering-can-accelerate-developer-velocity/feed/ 0
Easily monitor IBM i with updated Dynatrace extension https://www.dynatrace.com/news/blog/easily-monitor-ibm-i-with-updated-dynatrace-extension/ https://www.dynatrace.com/news/blog/easily-monitor-ibm-i-with-updated-dynatrace-extension/#respond Wed, 06 Mar 2024 16:55:35 +0000 https://www.dynatrace.com/news/?p=62883 IBM Monitoring Overview

Observability in large IT environments, such as the IBM i operating system, is critical for tracking system performance. To simplify this task, Dynatrace further enhanced an extension that enables effortless analysis of IBM i performance—without needing installations on your mainframe infrastructure. To collect performance data and provide insights into aspects like system health, job performance, and disk utilization, the extension links remotely to your IBM i system through Dynatrace ActiveGate.

The post Easily monitor IBM i with updated Dynatrace extension appeared first on Dynatrace news.

]]>
IBM Monitoring Overview

What is IBM i?

IBM i, formerly known as iSeries, is an operating system developed by IBM for its line of IBM i Power Systems servers. It is based on the IBM AS/400 system and is known for its reliability, scalability, and security features. IBM i is designed to integrate seamlessly with legacy and modern applications, allowing businesses to run critical workloads and applications. It also includes various built-in software components for database management, security, and application development.

Gaining knowledge about IBM i performance can be a challenging and pricey task. Some tools demand the installation of agents on those systems and provide complex, disconnected views. Additionally, certain tools require auxiliary services to gather performance data before it can be examined and queried.

Effortlessly analyze IBM i Performance with the new Dynatrace extension

Dynatrace has created a new version of its popular extension that is faster, offers better interactive pages, and includes more metrics, metadata, and analytics without having to install anything on your mainframe infrastructure.

Find out how your LPARs are doing overall, which jobs are consuming the most resources, and which users are holding up your output queues. Or perhaps you want to be notified when specific critical system or completion messages appear in your message queues.

Maybe you’d like to know when your jobs are in a waiting status, when an application is spawning too few or too many job instances, or when all the disks in your pool are being utilized evenly.

Get answers to all of these questions with the new Dynatrace IBM i extension.

How does the extension work?

The extension runs remotely from your Dynatrace ActiveGates and connects to your IBM i system. It then collects performance data using existing database services running on your system. The entire topology is built on Dynatrace based on the components it finds configured on your IBM i system, all acting cohesively to accelerate root cause and impact analysis.

Nothing is installed on your IBM i systems. It’s all monitored remotely!

More metrics, more data

We have improved performance KPI collection and added new metrics and entities, like:

  • System
  • Memory pools
  • Jobs
  • Job queues
  • ASP and disks
  • Output queues and spooling files
  • Message queues
  • Network
  • Subsystems

Each entity has more metadata to help you identify and understand its configuration.

All systems at a glance

A ready-to-use dashboard, straight out of the box, is the starting point, showing important information in a unified view.

Default dashboard for IBM I monitoring
Figure 1. Default dashboard for IBM I monitoring

The default dashboard provides an overview of all monitored systems and how many different entities are created by IBM i components.

Starting at this dashboard, you can drill down into any component to see how it performs on more detailed dashboards.

Get a health overview of each system

Monitor your system’s performance and detect unexpected events such as IPLs, CPU spikes, and exceeded total job limits. Identify users who are consuming the most spooling files.

System health overview
Figure 2. System health overview

Monitor your jobs and their instances

On IBM i, jobs are tasks that are triggered to run by the system and applications. It’s crucial to monitor the performance of these jobs, including their CPU usage, number of instances, and status. Additionally, it’s important to identify when a job is locked or hanging by observing various indicators.

Jobs dashboard
Figure 3. Jobs dashboard

Critical messages to your teams

Monitoring the main message queues can provide critical information about system failures, application failures, completion tasks, and any other message type you want to know about.

Messages overview
Figure 4. Messages overview

Monitor disks and disk pool utilization

One of the most important functions of your mainframe infrastructure is reading and writing data at high speeds while making it readily available. It’s critical to know the performance of your disks and ensure their optimum utilization to maintain balance in I/O operations. If you observe any disk running out of space more quickly than others—or underperforming—this type of analysis can help prevent disastrous outages or potential data loss.

Disk pool monitoring
Figure 5. Disk pool monitoring

As this extension evolves, your valued feedback is important to ensure its continuous improvement. Please visit our Product Ideas community forum, let your voice be heard, and give others a chance to upvote your ideas.

To learn how to deploy this extension, please visit Dynatrace Hub.

The post Easily monitor IBM i with updated Dynatrace extension appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/easily-monitor-ibm-i-with-updated-dynatrace-extension/feed/ 0
Kubernetes vs Docker: What’s the difference? https://www.dynatrace.com/news/blog/kubernetes-vs-docker/ https://www.dynatrace.com/news/blog/kubernetes-vs-docker/#respond Sat, 02 Mar 2024 12:59:14 +0000 https://www.dynatrace.com/news/?p=46533 Enhancing Kubernetes cluster management key to platform engineering success

If cloud-native technologies and containers are on your radar, you’ve likely encountered Docker and Kubernetes and might be wondering how they relate to each other. Is it Kubernetes vs Docker or Kubernetes and Docker—or both? What is the difference between Kubernetes and Docker? Docker is a suite of software development tools for creating, sharing and […]

The post Kubernetes vs Docker: What’s the difference? appeared first on Dynatrace news.

]]>
Enhancing Kubernetes cluster management key to platform engineering success

If cloud-native technologies and containers are on your radar, you’ve likely encountered Docker and Kubernetes and might be wondering how they relate to each other. Is it Kubernetes vs Docker or Kubernetes and Docker—or both?

What is the difference between Kubernetes and Docker?

Docker is a suite of software development tools for creating, sharing and running individual containers; Kubernetes is a system for operating containerized applications at scale.

Think of containers as standardized packaging for microservices with all the needed application code and dependencies inside. Creating these containers is the domain of Docker. A container can run anywhere, on a laptop, in the cloud, on local servers, and even on edge devices.

A modern application consists of many containers. Operating them in production is the job of Kubernetes. Since containers are easy to replicate, applications can auto-scale: expand or contract processing capacities to match user demands.

Docker and Kubernetes are mostly complementary technologies—Kubernetes and Docker. However, Docker also provides a system for operating containerized applications at scale, called Docker Swarm—Kubernetes vs Docker Swarm. Let’s unpack the ways Kubernetes and Docker complement each other and how they compete.

What is Docker?

Just as people use Xerox as shorthand for paper copies and say “Google” instead of internet search, Docker has become synonymous with containers. Docker is more than containers, though.

Docker is a suite of tools for developers to build, share, run and orchestrate containerized apps.

Docker container architecture
Docker container architecture

How does Docker work?

In Docker’s client-server architecture, the client talks to the daemon, which is responsible for building, running, and distributing Docker containers. While the Docker client and daemon can run on the same system, users can also connect a Docker client to a remote Docker daemon.

Developer tools for building container images

Docker Build creates a container image, the blueprint for a container, including everything needed to run an application – the application code, binaries, scripts, dependencies, configuration, environment variables, and so on. Docker Compose is a tool for defining and running multi-container applications. These tools integrate tightly with code repositories (such as GitHub) and continuous integration and continuous delivery (CI/CD) pipeline tools (such as Jenkins).

Sharing images

Docker Hub is a registry service provided by Docker for finding and sharing container images with your team or the public. Docker Hub is similar in functionality to GitHub.

Running containers

Docker Engine is a container runtime that runs in almost any environment: Mac and Windows PCs, Linux, and Windows servers, the cloud, and on edge devices. Docker Engine is built on top containerd, the leading open source container runtime, a project of the Cloud Native Computing Foundation (CNCF).

Built-in container orchestration

Docker Swarm manages a cluster of Docker Engines (typically on different nodes) called a swarm. Here the overlap with Kubernetes begins.

What is Kubernetes?

Kubernetes is an open source container orchestration platform for managing, automating, and scaling containerized applications. Kubernetes is the de facto standard for container orchestration because of its greater flexibility and capacity to scale, although Docker Swarm is also an orchestration tool.

Kubernetes architecture
A Kubernetes cluster is made up of nodes that run on containerized applications.

How does Kubernetes work?

A Kubernetes cluster is made up of nodes that run on containerized applications. Every cluster has at least one worker node. The worker node hosts the pods, while the control plane manages the worker nodes and pods in the cluster.

What is Kubernetes used for?

Organizations use Kubernetes to automate the deployment and management of containerized applications. Rather than individually managing each container in a cluster, a DevOps team can instead tell Kubernetes how to allocate the necessary resources in advance.

Where Kubernetes and the Docker suite intersect is at container orchestration. So when people talk about Kubernetes vs. Docker, what they really mean is Kubernetes vs. Docker Swarm.

For a deeper look into how to gain end-to-end observability into Kubernetes environments, tune into the on-demand webinar Harness the Power of Kubernetes Observability.

Benefits of using Kubernetes and Docker

When used together, Kubernetes and Docker containers can provide several benefits for organizations that want to deploy and manage containerized applications at scale.

Some of the key benefits of using include:

  • Scalability: Kubernetes can scale containerized applications up or down as needed, ensuring they always have the resources they need to perform optimally. This is helpful for applications that experience a boost in traffic or demand.
  • High availability: Kubernetes can ensure that containerized applications are highly available by automatically restarting containers that fail or are terminated. This can keep applications running smoothly and prevent downtime.
  • Portability: Docker containers are portable, meaning they can be easily moved from one environment to another. This makes it easy to deploy containerized applications across different infrastructures, such as on-premises servers, public cloud providers, or hybrid environments.
  • Security: Kubernetes can secure containerized applications by providing role-based access control, network isolation, and container image scanning. This can help to protect applications from unauthorized access, malicious attacks, and data breaches.
  • Ease of use: Kubernetes can automate the deployment, scaling, and management of containerized applications. This can save organizations time and resources, and it can also help to reduce the risk of human error.
  • Reduce costs: By automating the deployment and management of containerized applications, Kubernetes and Docker can help organizations reduce IT operations costs.
  • Improve agility: Kubernetes and Docker can help organizations to be more agile by making it easier to deploy new features and updates to applications.
  • Increase innovation: Kubernetes and Docker can help organizations to innovate more quickly by providing a platform that is easy to use and scalable.

Use cases for Kubernetes and Docker

When used in tandem, Kubernetes and Docker create a dynamic duo that unlocks a myriad of possibilities for seamless and scalable application deployment.

Here are a few use cases for using Kubernetes and Docker together:

  • Deploying and managing microservices applications: Microservices applications are made up of small, independent components that can be easily scaled and deployed. Each microservice can be containerized using Docker, and Kubernetes can manage the deployment and scaling of these services independently. This allows for better maintainability, scalability, and fault isolation.
  • Dynamic scaling: Together, Kubernetes and Docker enable dynamic scaling of applications. Kubernetes can automatically adjust the number of application instances based on demand. When traffic spikes, new containers can be spun up to handle the load, and when the load decreases, excess containers can be scaled down. This elasticity ensures efficient resource utilization and cost savings.
  • Running containerized applications on edge devices: Kubernetes can be used to run containerized applications on edge devices, ensuring that they’re always available and up to date. Docker has redefined how applications are packaged and isolated. Docker eliminates the “it works on my machine” dilemma by encapsulating an application and its dependencies within a standardized container. This consistency ensures the application runs the same way across development, testing, and production environments.
  • Continuous integration and continuous delivery (CI/CD): The combination of Docker and Kubernetes streamlines CI/CD pipelines. Docker images can be integrated into the CI/CD process, ensuring consistent testing and deployment. Kubernetes automates the deployment process, reducing manual intervention and accelerating the time to market for new features.
  • Cloud-native applications: Docker and Kubernetes are cloud-agnostic, making it easier to deploy applications across different cloud providers or hybrid environments. This flexibility allows organizations to choose the most suitable infrastructure while avoiding vendor lock-in.

What are the challenges of container orchestration?

Although Docker Swarm and Kubernetes both approach container orchestration a little differently, they face the same challenges. A modern application can consist of dozens to hundreds of containerized microservices that need to work together smoothly. They run on multiple host machines, called nodes. Connected nodes are known as a cluster.
Hold this thought for a minute and visualize all these containers and nodes in your mind. It becomes immediately clear there must be a number of mechanisms in place to coordinate such a distributed system. These mechanisms are often compared to a conductor directing an orchestra to perform elaborate symphonies and juicy operas for our enjoyment. Trust me, orchestrating containers is more like herding cats than working with disciplined musicians (some claim it’s like herding Schrödinger’s cats). Here are some of the tasks orchestration platforms are challenged to perform.

Container deployment

In the simplest terms, this means to retrieve a container image from the repository and deploy it on a node. However, an orchestration platform does much more than this: it enables automatic re-creation of failed containers, rolling deployments to avoid downtime for the end-users, as well as managing the entire container lifecycle.

Scaling

This is one of the most important tasks an orchestration platform performs. The “scheduler” determines the placement of new containers so compute resources are used most efficiently. Containers can be replicated or deleted on the fly to meet varying end-user traffic.

Networking

The containerized services need to find and talk to each other in a secure manner, which isn’t a trivial task given the dynamic nature of containers. In addition, some services, like the front-end, need to be exposed to end-users, and a load balancer is required to distribute traffic across multiple nodes.

Observability

An orchestration platform needs to expose data about its internal states and activities in the form of logs, events, metrics, or transaction traces. This is essential for operators to understand the health and behavior of the container infrastructure as well as the applications running in it.

Security

Security is a growing area of concern for managing containers. An orchestration platform has various mechanisms built in to prevent vulnerabilities such as secure container deployment pipelines, encrypted network traffic, secret stores and more. However, these mechanisms alone are not sufficient, but require a comprehensive DevSecOps approach.

With these challenges in mind, let’s take a closer look at the differences between Kubernetes and Docker Swarm.

Kubernetes vs Docker Swarm

Both Docker Swarm and Kubernetes are production-grade container orchestration platforms, although they have different strengths.
Docker Swarm, also referred to as Docker in swarm mode, is the easiest orchestrator to deploy and manage. It can be a good choice for an organization just getting started with using containers in production. Swarm solidly covers 80% of all use cases with 20% of Kubernetes’ complexity.

Docker Swarm architecture
Docker Swarm architecture

A swarm is made up of one or more nodes, which are physical or virtual machines running in Docker Engine.

Swarm seamlessly integrates with the rest of the Docker tool suite, such as Docker Compose and Docker CLI, providing a familiar user experience with a flat learning curve. As you would expect from a Docker tool, Swarm runs anywhere Docker does and it’s considered secure by default and easier to troubleshoot than Kubernetes.

Kubernetes, or K8s for short, is the orchestration platform of choice for 88% of organizations. Initially developed by Google, it’s now available in many distributions and widely supported by all public cloud vendors. Amazon Elastic Kubernetes Service, Microsoft Azure Kubernetes Service, and Google Kubernetes Platform each offer their own managed Kubernetes service. Other popular distributions include Red Hat OpenShift, Rancher/SUSE, VMWare Tanzu, IBM Cloud Kubernetes Services. Such broad support avoids vendor lock-in and allows DevOps teams to focus on their own product rather than struggling with infrastructure idiosyncrasies.

The true power of Kubernetes comes with its almost limitless scalability, configurability, and rich technology ecosystem including many open-source frameworks for monitoring, management, and security.

Kubernetes vs. Docker Swarm

Docker and Kubernetes: Better together

Simply put, the Docker suite and Kubernetes are technologies with different scopes. You can use Docker without Kubernetes and vice versa, however they work well together.

From the perspective of a software development cycle, Docker’s home turf is development. This includes configuring, building, and distributing containers using CI/CD pipelines and DockerHub as an image registry. On the other hand, Kubernetes shines in operations, allowing you to use your existing Docker containers while tackling the complexities of deployment, networking, scaling, and monitoring.

Although Docker Swarm is an alternative in this domain, Kubernetes is the best choice when it comes to orchestrating large distributed applications with hundreds of connected microservices including databases, secrets, and external dependencies.

How does advanced observability benefit Kubernetes and Docker Swarm?

Whether you’re using Kubernetes or Docker Swarm, or both, managing clusters at scale comes with unique challenges, particularly when it comes to observability. Application teams and Kubernetes/Swarm platform operators alike depend on detailed monitoring data. Here are some examples.

How advanced observability benefits Kubernetes and Docker Swarm

Kubernetes provides some very basic monitoring capabilities, like event logs and CPU loads for example. However, there’s a growing number of open-standard and open-source technologies available to augment Kubernetes’ built-in features. Some frequently used observability tools include: Promtail, Fluentbit and Fluentd for logs; Prometheus for metrics; and OpenTelemetry for traces, to name a few.

Dynatrace integrates with all these tools and more, and adds its own high-fidelity data to create a single real-time entity model. This unique capability enables Dynatrace to provide advanced analytics, AI-powered root-cause-analysis and intelligent automation, providing application teams and platform operators a unified view on the full technology stack.

To learn more about how Dynatrace automatically ingests and leverages data from open-source tools like Prometheus, Fluentbit, and OpenTelemetry, join us for the on-demand Performance Clinic Kubernetes observability for SREs with Dynatrace.

The post Kubernetes vs Docker: What’s the difference? appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/kubernetes-vs-docker/feed/ 0
AI techniques enhance and accelerate exploratory data analytics https://www.dynatrace.com/news/blog/ai-techniques-accelerate-exploratory-data-analytics/ https://www.dynatrace.com/news/blog/ai-techniques-accelerate-exploratory-data-analytics/#respond Wed, 28 Feb 2024 19:49:01 +0000 https://www.dynatrace.com/news/?p=62690 Causal AI use cases for modern observability; exploratory data analytics

To make exploratory data analytics even easier, organizations are using more AI techniques to make sense of data from their cloud environments. With Dynatrace Grail dashboards, Notebooks, and CoPilot generative AI, getting instant answers has never been easier.

The post AI techniques enhance and accelerate exploratory data analytics appeared first on Dynatrace news.

]]>
Causal AI use cases for modern observability; exploratory data analytics

In a digital-first world, site reliability engineers and IT data analysts face numerous challenges with data quality and reliability in their quest for cloud control. Increasingly, organizations seek to address these problems using AI techniques as part of their exploratory data analytics practices.

Exploratory data analytics is an analysis method that uses visualizations, including graphs and charts, to help IT teams investigate emerging data trends and circumvent issues, such as unexpected traffic spikes or performance degradations.

Challenges to exploratory data analytics

Among the challenges analysts face are multiple heterogeneous data sources, noisy or incomplete data, insufficient causal reasoning (faulty connections between event cause and effect), and untrustworthy AI, according to an article from the Columbia University Data Science Institute.

Another hurdle is mistaking easy patterns as effective analysis, according to an article in the Harvard Data Science Review. Techniques analysts use to emphasize patterns, such as aggregating data by default, can cause them to overlook variation and uncertainty in their data, so they can draw conclusions that the data don’t fully support.

AI techniques provide a solution

A first line of attack is selecting the right analytics tool, which can help teams detect meaningful patterns in real-time data, integrate data from multiple sources, and render data visualizations.

To that end, in 2022, Dynatrace released Grail, the auto-indexing, schema-on-read data lakehouse, along with Notebooks and Dashboards. This expansion of the Dynatrace platform builds on its foundation, which uses topology mapping and causal AI, an AI technique based on fault-tree analysis, to pinpoint root causes.

The next challenge is harnessing additional AI techniques to make exploratory data analytics even easier. Dynatrace Grail, Notebooks, Dashboards, and multiple AI techniques combine to provide analysts with instant insights, as demonstrated by Thomas Ziegelbecker, a senior product manager at Dynatrace, and his colleagues Peter Zahrer, principal product manager, and Gabriele Hasson-Birkenmayer, senior product manager at the recent Perform 2024 conference.

Three steps in exploratory data analytics: Discover, browse, explore

Grail captures heterogeneous data from across the network in one place while retaining its context and semantic details, which eliminates the limitations of traditional databases. From this unified, semantically rich data resource, analysts can explore data on the fly and share findings using Notebooks and Dashboards.

With Notebooks, analysts can “explore their data, use ad-hoc analysis, and work with data in a sequential fashion to refine, refine, refine,” Ziegelbecker said. “With Dashboards, you can observe [the data you’re interested in] over time,” covering common monitoring use cases such as system health.

Exploratory data analytics phases: Discover, browse, explore

“When you work with data, it comes down to three steps: Discover, browse, and explore,” Zielgelbecker said. “Start by asking yourself what’s there, whether it’s logs, metrics, or traces. Once you double down on the type, you want to figure out, or browse, which metric is relevant. Then when you have the metric, you want to explore it—what splits are available, how can I break it down, how can I aggregate it, and so on.”

Discover data using global search on the new ‘explore’ section and tile type in Notebooks and Dashboards

From a dashboard, an analyst can investigate an issue, such as a 25% error rate from a key Kubernetes cluster. According to Ziegelbecker, the discovery phase of exploratory data analytics can start in two ways: doing a global search or using the “explore data” interface.

AI techniques used to explore Kubernetes errors in logs

  1. Discovery using global search. Analysts can easily navigate to any entities (hosts, applications, processes, Kubernetes nodes, and so on) or metrics in their environment using global search. Users can trigger the global search from any context with CTRL/CMD +K. Type > to see a list of all available search categories. Select the relevant category (such as “Metrics”) and type the search term. This provides a quick way to navigate data or start a DQL query.
  2. Discovery using the “explore data” interface. For those who think visually, Dynatrace provides an interface to explore data within Dashboards and Notebooks. Using step navigation and drop-down menus, this simplified UI approach streamlines query building for essential tasks such as filtering, aggregating, and sorting. This method also enables users to aggregate, filter, sort, limit, or split metrics.

Browse data: Advanced exploration using DQL

Once analysts discover the data they’re interested in, the next step is to refine the search and share results.

Dynatrace Notebooks is an interactive data exploration interface that enables users to collaborate using code, text, and rich media to build, evaluate, and share insights.

“[Notebooks] is purposely built to focus on data analytics,” Zahrer said. “We use Dashboards to monitor and present; we use Notebooks to work with the data and, at the same time, document what we just did.”

Using DQL, users can query any data stored in Grail, such as metrics, logs, events, and time series, in context of any entity (hosts, processes, applications) in the monitored environment, provided by Dynatrace Smartscape.

Using Notebooks, an analyst can extend the query started in the discovery phase to further troubleshoot an issue, such as a 25% error rate in a Kubernetes cluster. “I’m interested in seeing whether the [Kubernetes] log errors we’re seeing on our dashboard are related to a misconfiguration on the Kubernetes side,” Zahrer said.

To relate logs and metrics using Grail, analysts can use DQL within the Notebooks app to further explore the problematic Kubernetes logs. DQL prompts help analysts filter on enriched data from Dynatrace OneAgent—a single file that automatically discovers all the processes running on a host and auto-instruments application pages.

By enriching data with the topological context of Smartscape through OneAgent and keeping data consistent, Zahrer said, Dynatrace helps analysts do causation-based cross-correlation to see how events relate and to pass on details, such as Kubernetes nodes involved in errors, to further refine investigation and analysis.

Explore data using Davis CoPilot—a generative AI technique for advanced analytics of logs and metrics

As an alternative to identifying and exploring data, analysts can also use Davis CoPilot to achieve the same result.

“Davis CoPilot is a generative AI that works in concert with our predictive AI and with our causal AI,” Hasson-Birkenmayer said. The blend of these three AI techniques—predictive, causal, and generative AI—make up the Davis composite, or hypermodal, AI approach.

Davis CoPilot generative AI helps analysts get started with—and get more proficient using—DQL. Instead of configuring tiles or sections, or instead of writing DQL, analysts can explore data with a natural language prompt. For example, “you can ask Davis CoPilot to ‘summarize all logs by status,’” she explained. “In the background, Davis converts the prompt into DQL and auto-executes it. So you can go straight from a conversational prompt using your natural language directly into insights.”

Because the Davis CoPilot integration in Notebooks is a DQL assistant, it runs the query so analysts can see if the results are what they intended without having to first review and validate the DQL syntax themselves. If users need to refine the results, they can do so either by refining the natural language input or by creating a new DQL section and refining the details directly.

Analysts can also customize data visualizations according to their needs. “In some cases, I am actually choosing and selecting a particular visualization, and in some cases [such as with certain metrics], it gets done automatically,” Hasson-Birkenmayer said.

Exploratory data analytics enhanced by AI techniques

Exploratory data analytics enhanced by AI techniques

Starting from a dashboard or notebook using the “Explore data” interface to cover simple data exploration routines, analysts can advance their inquiries in Grail using DQL. With a combination of AI techniques—generative AI using data verified by predictive and causal AI—Davis CoPilot enables analysts to further refine complex explorations using natural language queries.

Learn more about how Grail, Dashboards, Notebooks, and Davis CoPilot work together to speed up and refine exploratory data analytics in the on-demand session from Perform 2024, Your analytics superpower: Empowering teams to gain instant insights with Dynatrace.

The post AI techniques enhance and accelerate exploratory data analytics appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/ai-techniques-accelerate-exploratory-data-analytics/feed/ 0