Lately, I’ve been spending a lot of time talking with enterprise automation leaders about how they monitor their mission-critical business processes. Whether it’s overnight trading settlements or daily banking reconciliations, these are the workflows that keep the lights on. When they fail, it’s an immediate, highly visible crisis.
In these conversations, I keep seeing a very specific pattern emerge. Desperate to get better visibility across their complex, multi-platform environments, application and IT operations teams often take matters into their own hands. They start building custom dashboards, or they write scripts to pipe automation data into tools like Datadog, Splunk, or Grafana.
On the surface, this custom work makes sense. But what begins as a simple script or a well-intentioned dashboard almost inevitably morphs into a massive, resource-draining software product of its own.
It's a "do-it-yourself" trap, and it introduces some serious hidden risks to the business.
Field notes on the "do-it-yourself" trap
Why should organizations move from custom dashboards to purpose-built workload analytics and intelligence? When we look under the hood of homegrown monitoring tools, we usually find a few glaring issues that actually make operations harder, not easier.
The scheduler performance drain
To feed these custom dashboards, teams often write scripts that aggressively poll the automation engine for status updates. At one major financial services firm, application teams were running roughly 100,000 queries an hour against their scheduler just to monitor the status of individual jobs. This kind of primitive, brute-force polling creates massive resource bottlenecks, bogging down the actual automation engine that teams are trying to monitor.
The “automation monitoring automation” paradox
Another global bank built a custom system that used automation jobs to scrape the scheduler’s alert tables and feed them into their central enterprise paging system. It sounds like a clever workaround, until something breaks. In one instance, the alarm job failed. When it was finally restarted 12 hours later, it spammed the paging system with tens of thousands of delayed alerts, literally taking down the enterprise communication infrastructure with a flood of false positives.
The crushing, ever-expanding maintenance burden
What are the hidden costs of building custom workload automation monitoring tools? Typically, we have seen the following lifecycle:
-
Create a simple, custom application to monitor a specific set of jobs against a basic set of criteria, such as hard-wired response thresholds.
-
Instrument critical jobs with pre- and post-execution “probes.” (Ironically, these probes can introduce latency, complexity, and additional points of failure.)
-
Create new processes and data pipelines to collect and store the probe data.
-
Add features to the simple application to “monitor” incoming data from critical endpoints in near real time.
-
Add capabilities to this increasingly complex application, for example to enable alerting or notification or provide integrations with existing notification systems.
-
Add some reporting capabilities to the ever-growing application.
-
Try to add root cause analysis, backtracing, and triage capabilities to this increasingly substantive custom software.
-
Try to expand the now highly complex application to enable it to predict future events.
By around step 4 or 5, the application has taken on a life of its own, requiring a dedicated, full-time team to make platform updates, manage version currency, address security vulnerabilities, and simply keep up with users’ feature requests. Each step along this path becomes exponentially more difficult and expensive.
Over time, this can easily run into millions of dollars annually in maintenance burden, diverting scarce resources away from building revenue-generating business applications. Further, most of our largest customers find that around step 7, making additional progress presents a monumental challenge. It is at this point that the true cost of a DIY approach—and the value of Broadcom’s Automation Analytics & Intelligence (AAI)—becomes crystal clear.
Then there is the fact that many of these applications have been built in silos. Often, one scheduling team (for example, a group in the mainframe domain) builds their own tools with little or no awareness of the tools that another scheduling team (say, in the distributed arena) has already built, and vice-versa.
To add insult to injury, we finally arrive at the intersection of the platforms, where an issue on one side of the house sets a fire in the other side. Queue up the usual finger pointing, hand waving, and hair-on-fire scrambling to respond to emergencies.
The 5 a.m. fire drill
Perhaps the biggest issue is this: Even with these expensive custom tools, teams are still flying blind. Dashboards like the "the status ring" only turn red after an SLA has already been breached.
Because these DIY tools lack dynamic critical path and dependency mapping, a red milestone triggers a 5 a.m. fire drill. Suddenly, you have half a dozen senior operations people on a bridge call, trying to rapidly sift through tens of thousands of jobs manually to find the root cause. (Given all this, it’s no wonder recent industry research shows that 92% of organizations require more than an hour just to identify the root cause of an orchestration issue). This begs the question, how can enterprise IT teams prevent workload failures and improve SLA adherence?
Moving up the maturity curve
When I see these patterns, it reinforces that there’s a maturity curve in automation observability, and many organizations are in the early stages of that evolution.
Most teams using homegrown dashboards are stuck at stage 1: This leaves them reacting to binary alerts (that is, "Did it fail?") after the damage is already done.
To actually protect the business, you have to get to stage 2. This requires moving from a focus on "Did it fail?" to answering this question: "Will it be late?" This requires proactive, predictive SLA management and dynamic critical path analysis. Static BI tools and custom dashboards simply cannot perform the complex mathematics these capabilities require.
Right now, the most exciting conversations I'm having are with customers who are pushing even further, into stage 3 and the realm of financial intelligence and workload FinOps. (See how AAI v26 enables this move by introducing a robust financial intelligence model that enables organizations to calculate total cost of ownership (TCO) for their automation practices.)
The overhead challenge (and why FinOps matters)
Here is the other thing I hear constantly from IT leaders: Workload automation is treated like an "all-you-can-eat" buffet.
Right now, central IT absorbs 100% of the software licensing, mainframe MIPS, and cloud computing bills. Meanwhile, line-of-business users just submit requests and consume the automation as a free, unlimited utility. When IT presents its annual budget, workload automation appears as an opaque cost sink.
You can't manage what you can't measure. Organizations need a way to associate technical job runs with actual business outcomes.
This is where a purpose-built intelligence layer like AAI changes the game. Instead of requiring you to build custom scripts, AAI provides a non-intrusive way to calculate the total cost of ownership (TCO) of your workloads. It uses rule-based mapping to automatically associate tens of thousands of jobs directly to specific cost centers and business units—no manual tagging required.
Suddenly, you can generate defensible showback scorecards that pair the cost of a service with SLA reliability. IT shifts from simply being an infrastructure cost center to a transparent, strategic service provider.
Simplify and predict
You don't need to rip and replace your underlying schedulers to gain visibility, and you certainly don't need to dedicate a team of developers to maintaining a homegrown dashboard that still leaves you scrambling at 5 a.m.
We continue to see organizations succeeding with AAI. For example, a major commercial bank employed AAI to achieve these objectives:
-
Unify SLA management.
-
Replace their expensive custom-built repositories.
-
Achieve significant reduction in their legacy database costs.
By standardizing on a purpose-built intelligence and FinOps layer, you can protect your business SLAs, eliminate alert noise, and definitively prove the exact ROI of your automation practice.
Ready to move beyond the dashboards that leave you blind-sided by SLA breaches and saddled with spiraling costs? Explore how Broadcom’s Automation Analytics & Intelligence (AAI) can transform your workload orchestration.
Frequently asked questions
Q: Why do homegrown workload automation monitoring approaches bog down operations teams?
A: Custom dashboards demand significant time and cost just to get started, and the costs only keep growing. Over time, these applications continue to need updates and refinements, ultimately requiring dedicated teams to maintain.
Q: What is the primary difference between stage 1 and stage 2 automation observability?
A: When operating in stage 1, teams are left reacting after an SLA breach or failure has already occurred. On the other hand, in stage 2, teams use proactive, predictive SLA management and critical path analysis to forecast delays, so they can be mitigated before they have an impact on operations.
Q: How does workload FinOps help organizations manage automation costs?
A: Workload FinOps uses rule-based mapping to link technical job runs directly to specific cost centers and business units. In this way, operations teams can move from an opaque cost center to a strategic service provider.
Q: Do organizations need to replace their existing schedulers to use Automation Analytics & Intelligence (AAI)?
A: No, organizations can keep using their existing schedulers. AAI functions as a non-intrusive intelligence layer that unifies visibility across complex, multi-platform scheduling environments.