When Millions Tune In at Once, Observability Has to Keep Up
The hardest and most stressful moment for a streaming platform during a live sports event may be a few minutes before the opening whistle. That moment can turn a relatively ordinary streaming workload into something entirely different. Viewership can jump by millions in minutes, traffic can move rapidly between content delivery networks (CDNs), and infrastructure teams have to make decisions while the growing audience is tuning in.
For the people responsible for keeping a streaming service running, there is little room for a delayed answer. If something goes wrong, they need to know the origin of the problem, whether it’s which CDN is impacted, the internet service provider, a particular network, a player, or somewhere else in the delivery chain. And they need that answer at that moment, while the event is happening.
That means the observability platform itself has to scale with the event.
“Streaming providers and broadcasters are being forced to rethink their observability posture,” says Coralie Seuzaret, Senior Solutions Engineer at Hydrolix, a real-time data platform. “Even if you have an observability platform that works perfectly well for business-as-usual traffic, a major live event changes everything.”
The challenge is not simply that there is more data. There are more data sources, more stakeholders, more CDNs, and more decisions happening simultaneously. A multi-CDN streaming operation may need to correlate CDN access logs with player-quality data, origin logs and other telemetry, to decide when and where to redirect traffic.
During an event like the 2026 World Cup, when Hydrolix was used for observability by a major UK broadcaster, those loads can be greater than normal. And the people operating the infrastructure can’t wait until the next morning to investigate. For broadcasters with rights to games airing around the clock, operations teams can spend weeks working around the clock, including weekends. An issue can occur at any time, and the consequences extend beyond technical performance. A degraded stream can affect audience experience, advertising revenue, and a broadcaster's reputation.
Historically, one way to make observability systems more manageable and affordable has been to sample logs. But that means when an incident happens, operators piece together evidence with a partial data set. The discarded data could be the evidence that leads to the root cause. Without it, operators may go down the wrong path and remediation may take longer.
“When there is that much at stake, you can't afford to focus on the wrong 10% of the data,” Seuzaret says.
Hydrolix has encountered situations where sampled data suggested that everything was working normally. For example, one broadcaster (who was not part of the World Cup) was receiving complaints from users that they were seeing an end-of-stream message pop up in the middle of a game. But the broadcaster could only see successful requests. Hydrolix showed all the data unsampled, revealing false successful requests, picking up on the error codes that were returned, and the CDN misconfiguration that led to it.
The same issue becomes increasingly important as streaming companies begin applying AI to their operational data. An AI agent can only reason about the information it receives. If an agent is operating on a sample that represents 10% of the underlying traffic, it can miss the events that matter, drift, and potentially make the wrong decisions.
“Major live event streaming is one of those moments where you shouldn't take that risk,” says Seuzaret. “If you build agent workflows on data that represents only part of your volume, you can miss things. You can invest in and develop a great agent and still end up with a bad decision due to the lack of context and data fidelity.”
That makes real-time, full-fidelity data a central requirement, not an optional feature.
Hydrolix's work with the 2026 World Cup provided a real-world test of that approach. The broadcaster's streaming service ultimately handled 51 games across 30 days, with more than 9 million concurrent viewers at peak. Across the tournament, the peak CDN throughput reached 28.5 TB per second.
Behind that audience was an enormous stream of telemetry. Hydrolix ingested approximately 400 billion requests, representing almost 600 TB of unsampled data over the tournament. During the match against Croatia, ingestion peaked at more than 2.3 million rows per second.
The platform compressed the data by more than 90%, helping make retention and analysis of that volume economically practical. But ingestion capacity was only half of the equation. The data had to be usable while the match was underway.
Across the tournament, the broadcaster ran approximately 10 million queries against the data, which was roughly 190,000 queries per game. The average query execution time was under two seconds. The system was not simply collecting enormous quantities of telemetry and analyzing it after the fact. It was sustaining more than two million rows of incoming data per second while returning queries in less than two seconds, without sampling.
“Real time was non-negotiable,” says Seuzaret. “They needed to see what was happening while it was happening. We had to ingest the volume, scale quickly and still give them answers in the timeframe that mattered.”
And the scale changed quickly. During the opening ceremony, viewership climbed to almost 4 million concurrent viewers. That is the kind of traffic surge that can expose a weakness not only in a streaming platform, but in the observability infrastructure monitoring it.
“If your observability platform takes hours to scale, it doesn't help you during that ten-minute ramp,” Seuzaret says. “The platform has to be efficient at scaling up without creating an enormous cost problem for the customer.”
The broadcaster was operating across multiple CDNs, including Akamai, Fastly and CloudFront, which created another challenge. Every provider produces telemetry in their own format, and operators need to understand the combined picture rather than investigate each CDN in isolation.
For the World Cup, Hydrolix standardized the data so that CDN logs could be brought into a common schema and viewed together at once. Instead of maintaining separate analytical views for each CDN, the broadcaster could work from a unified view of the complete delivery environment.
That visibility could then be combined with other streaming telemetry, including Common Media Client Data (CMCD), giving operators another way to connect infrastructure behavior with the viewer experience. The system also supported more specific views when teams needed them, by channel, autonomous system number (ASN) or other dimensions.
The underlying deployment was designed to move quickly. For the broadcaster, much of the cluster deployment was automated, with the environment ready within hours. Data pipelines and schemas were prepared to accept the different CDN sources, allowing the customer to effectively plug into an already-prepared observability environment.
That speed matters because not every broadcaster starts preparing months in advance. Hydrolix worked with streaming providers across North America, Europe and India ahead of the World Cup. While some organizations had been preparing for months, others were looking for a solution only weeks before the tournament. The common factor was that the event itself could not be delayed while an observability architecture was designed.
Hydrolix managed the observability environment and monitored both the data and the underlying cloud infrastructure during the tournament, which meant the broadcaster didn’t have to spend the game thinking about the platform.
“The provider wants to focus on the game,” she says. “They don't want to be asking whether the observability platform can handle the scale. They need to know that we're watching it, that it will just work and that we'll be there if something changes.”
During the tournament, the broadcaster could monitor how much traffic was moving through each CDN and could also use the data to understand traffic patterns after individual matches.
The World Cup also illustrated how streaming observability is expanding beyond traditional uptime and troubleshooting. For example, it’s expanding into anti-piracy. Full-fidelity CDN data can provide clues about suspicious behavior, including reused tokens and unusual access patterns. Even when a streaming platform does not yet use token-based controls, patterns in CDN data can help identify potential piracy.
Sampling data makes that substantially harder. A sophisticated pirate may generate a small number of requests that look insignificant in isolation but become significant when viewed alongside the complete traffic record.
AI is another emerging use case. Hydrolix has built an MCP server that allows users to query Hydrolix data using natural language. That opens the possibility of generating post-game reports or investigating operational questions conversationally.
The quality of the underlying data remains critical. Giving an AI system access to the complete dataset provides more context and reduces the risk that it will reach a conclusion based on an incomplete sample. Full fidelity data is the AI’s context. The more data it’s fed, the more accurate conclusions it can make.
The World Cup experience points to several broader lessons for streaming providers preparing for major live events. First, plan for the traffic spike, not the average. An observability architecture that performs well during normal traffic may behave very differently when viewers arrive by the millions in minutes.
Second, treat observability capacity as part of event capacity. It is not enough for the streaming platform to scale if the systems monitoring it cannot keep up.
Third, avoid assuming that sampling is automatically good enough. AI and agentic are creating the need for full fidelity data, and with high compression, the costs don’t have to be astronomical. Also, during a high-stakes event, the missing data may contain the problem you're trying to find.
Fourth, unify the data. Multi-CDN delivery means multiple providers, formats, and operational stakeholders. A common data model can turn several fragmented views into one picture of the delivery environment.
Fifth, make the observability layer operationally invisible. The best time for an infrastructure team to discover a limitation in its monitoring system is not during the championship match. Automated scaling, capacity planning, and vendor support can take some of that burden away.
And finally, design for what comes after the event. The same data used to troubleshoot a live match can support CDN optimization, anti-piracy work, content steering, post-event analysis, and increasingly AI-driven operations.
Major live events are not getting simpler. Viewership can spike faster, delivery architectures can involve more CDNs and platforms, and the amount of telemetry generated by each viewer continues to grow. The lesson from the World Cup is therefore less about one tournament than about a change in the economics and architecture of streaming. When millions of people are watching simultaneously, observability has to operate at the same scale, speed, and fidelity as the stream itself.
Related Articles
Eluvio CEO and co-founder Michelle Munson says the streaming industry is entering a new phase where ultra-low-latency performance is no longer a niche requirement but a fundamental expectation for global sports. After a summer of high-pressure testing during the FIFA World Cup, she argues that Eluvio's Content Fabric has demonstrated a level of consistency that conventional cloud/CDN architectures simply cannot match.
03 Sep 2026
Gaming is built as an always-on system designed to maximize session frequency, retention, and lifetime engagement. Live sports, on the other hand, are still optimized for scheduled viewing windows where engagement peaks during the match and drops off immediately once it's over. For streaming product leaders and sports media executives responsible for retention and average revenue per user (ARPU), the opportunity isn't to change the live event. It's to capture the audience before and after it.
16 Apr 2026
AWS is the latest vendor to offer a solution to reformatting and syndicating vertical video for live sports as rights holders look to capitalize on the mobile first boom. Developed over 18 months with beta customers NBC Sports and Fox Sports, the ambition goes beyond reformatting highlights for vertical viewing but potentially live streaming whole games in the format.
24 Feb 2026