AFI OPS symbol

AFI OPS

Case Studies

About

Contact Us

← Back to case studies

Oves Enterprise

Cloudera & Spark Optimization: Stabilizing a Petabyte-Scale Data Platform

Ongoing
Real-time & Industrial Data Platforms
Product Engineering
Platform & Cloud Engineering

Platform scale

petabytes

Spark job cadence

unreliable

daily

Cluster stability

corrupting

restored

Oves Enterprise builds analytics and consulting products for enterprises, combining industry expertise with large-scale data processing.

What changed

Before

Cluster

Cloudera

corruption in the distribution

Jobs

Spark

long, failing runs

Availability

Critical services down

After

Cluster

Cloudera

root causes fixed

New nodes

Jobs

Spark

algorithms reworked

Cadence

Daily runs

petabyte scale

A corrupting Cloudera cluster stabilised, then scaled out and tuned until the jobs could run every day.

What we built

We partnered with the client to stabilize and optimize an enterprise navigation application running on Cloudera and Spark, hosted entirely on-premise in a local data center. The platform was handling petabyte-scale workloads and had accumulated a range of systemic issues that were impacting performance, availability, and long-term maintainability.

Challenges Addressed

Corrupt Cloudera installation causing instability and service failures.

Growing data volumes (petabytes) requiring additional nodes for scalability.

Spark jobs running inefficiently, leading to unacceptably long processing times.

Hadoop configuration not fine-tuned for replication and resilience.

Vendor dependency on Cloudera limiting long-term flexibility.

Minimal monitoring and alerting, making it difficult to ensure system availability.

Results

Identified and fixed root causes of corruption in the Cloudera distribution.

Restored stability and availability of critical data services.

Added new nodes to handle petabyte-scale workloads.

Improved Spark job algorithms, significantly reducing execution times and enabling daily job runs.

Fine-tuned resource allocation across YARN for optimized job scheduling.

Tuned replication factors, I/O operations, and failover configurations.

Migrated from Cloudera to a vanilla Hadoop distribution with HA master, eliminating vendor lock-in.

Set up Prometheus and Grafana for real-time monitoring and alerting.

Stack

Java
Jenkins
Apache Spark
Bash
PostgreSQL
Python
C#
Scala

Engagement: Ongoing · Real-time & Industrial Data Platforms, Product Engineering, Platform & Cloud Engineering

Next step

Let's look at your platform together.

45 minutes · an engineer, not a sales rep · no obligation