About Adobe
San Jose, California, based Adobe empowers everyone to create through industry-leading platforms and tools that unleash creativity, productivity, and personalized customer experiences.
www.adobe.com
Learn how ScaleOps with Amazon Web Services (AWS) helped Adobe accelerate AI development and eliminate waste while reducing time spent on resource management and ensuring consistent performance
Overview
Adobe is one of the world’s leading tech companies that has transformed the tools and manner of producing, publishing, and sharing digital media. As the company sought to increase its AI profile through its cloud-native developer platform, it recognized that it was running hundreds of clusters across a large number of business units leading to more than one million CPUs and thousands of GPUs to power its AI development and products. Adobe had to get a handle on eliminating waste even as it reduced time spent on resource management. Two objectives at seeming odds with each other. That is, until it integrated the ScaleOps Autonomous Cloud and AI Infrastructure Resource Management platform.
Too many clusters, not enough control
Adobe, one of the most recognizable brands in the world since 1982, has revolutionized digital media and publishing. The company employs about 30,000 people globally and is a leading player in AI with products such as Adobe Firefly, which has been used to generate more than 9 billion images across the software suite. As part of its AI program, Adobe Ethos (Core-K8 Platform) is the company’s internal, cloud-native developer platform.
As such, it runs hundreds of Amazon Elastic Kubernetes Service (Amazon EKS) clusters across numerous business units including Sensei, Firefly, Adobe Experience Manager, and Adobe Experience Platform with more than one million CPUs and thousands of GPUs powering AI products.
To meet business needs, Adobe concluded that it needed to accelerate its AI development and eliminate waste while reducing time spent on resource management. Its in-house rightsizing tool had low adoption due to performance issues and developers. Developers preferred over-provisioning to ensure availability. Therefore, Adobe was spinning up thousands of GPUs that had a very low utilization, creating resource bottlenecks. This is nothing new as it is relatively common for engineering teams to over-provision Kubernetes to stay safe, especially when rightsizing hundreds of services manually makes it impossible to remain current. This creates wasted capacity, inconsistent performance under load, and reliability incidents such as maxed out autoscaling thresholds and Horizontal Pod Autoscaling (HPA) limits set too low. Ergo, the trade-off between performance against efficiency.
ScaleOps: A continuous automatic optimization solution
Knowing it needed to bring in a partner, Adobe engaged in a period of due diligence before recognizing the clear advantages of ScaleOps.
Most tools only surface recommendations leaving engineers to act on them. ScaleOps applies optimization automatically and continuously at the workload and node level to adapt to real-time demand. The ScaleOps solution enhances native AWS tooling rather than replacing it so that ScaleOps sits between Karpenter and Amazon EKS to improve both. ScaleOps also works alongside HPA instead of conflicting with it. The result is automation that holds up in production, not static reports.
The ScaleOps solution for Adobe included Continuous Pod Rightsizing, where the solution sets requests and limits in real-time based on actual usage. Smart Bin Packing removes consolidation restraints by intelligently placing tools onto fewer nodes that leverage Karpenter. GPU Optimization raises GPU utilization and reclaims idle capacity on the same hardware by rightsizing and running multiple pods on a single GPU. Karpenter Node Optimization identifies and selects ideal node types for Karpenter to provision. In all, the ScaleOps platform continuously automates resource management to deliver performance, stability, and efficiency without manual tuning. This means real business value and competitive advantage.
Underscoring this utility is the ScaleOps-AWS partnership. AWS runs the largest base of production Kubernetes through Amazon EKS, which is exactly where ScaleOps delivers value. Partnering with AWS lets ScaleOps meet customers where they already build, integrate with native services like Karpenter, and co-sell with AWS teams to help customers run Amazon EKS more efficiently and reliably while scaling workloads with confidence. ScaleOps users on AWS consolidate optimization into one automated platform, free up engineering time, and can reinvest reclaimed capacity in AWS. And procurement via AWS Marketplace and AWS co-sell support shortens the time to value.
ScaleOps and AWS help Adobe achieve significant time and resource savings
During the proof of concept (POC) phase of the engagement, ScaleOps on AWS automated 240 of 516 Amazon EKS clusters. This projected out to yearly savings of $7 million based just on the POC. With full deployment, Adobe projects to save $30 million per year. All while maximizing performance, freeing engineers to focus on high-value tasks, and reducing time and resource waste. Additionally, the success of the ScaleOps-Adobe engagement led Adobe to be nominated for the AWS re:Invent keynote, a premier cloud and AI conference. Adobe is also renewing and expanding its ScaleOps deployment to achieve even more efficiency and savings.
Engineering teams can stop manually managing infrastructure and ship faster. Clusters self-tune as workloads change, performance stays consistent through traffic spikes, and reclaimed capacity is reinvested into more workloads on Amazon EKS. Optimization becomes a continuous, automated practice rather than a periodic project. Further, the AWS Marketplace streamlines procurement by allowing organizations to buy through existing AWS agreements while the ScaleOps-AWS partnership allows for deep, native integration points with Amazon EKS and Karpenter. AWS co-sell and Solutions Architect enablement extend ScaleOps’ reach while its growing GPU capacity on Amazon EKS helps organizations use accelerated compute more efficiently. Regardless of the organization’s size, ScaleOps on AWS offers a powerful tool to conserve resources while freeing engineering capacity.
