Case study

Real-Time Content Discovery on AWS OpenSearch

How SquareShift built a serverless AWS OpenSearch pipeline that powers real-time content discovery and behavioral analytics for a global AI-powered learning platform.

Book a session
100%Serverless AWS architecture built on Glue and OpenSearch
30GB of new content ingested daily
2,000+Search queries served per second at peak
6Terabyte cluster powering real-time content discovery

A global AI-powered learning platform used by enterprises and governments, processing more than a billion user events every day.

That real-time behavioral data drives the personalized learning and engagement experience the platform is built around.

Impact

Serverless AWS architecture built on Glue and OpenSearch. GB of new content ingested daily. Search queries served per second at peak.

Key services
PePlatform & Software Engineering
DaData & Analytics
ClCloud Modernization
Industry

Education

Key technologies / platforms

AWS · OpenSearch · AWS Glue · PostgreSQL · MySQL · InfluxDB

The engagement

How SquareShift delivered it.

The challenge

The platform needed real-time behavioral analytics with minimal latency, at a scale of over a billion user events a day. Ingestion had to pull from multiple sources at once — including legacy data still sitting in PostgreSQL and MySQL — and deduplicate it without slowing anything down.

What we delivered

SquareShift migrated more than a terabyte of historical data to AWS and built a fully serverless architecture on AWS Glue and OpenSearch. The pipeline now ingests roughly 30GB of new content every day, running through a 6TB OpenSearch cluster tuned to serve over 2,000 queries per second.

SquareShift consolidated the multiple source systems into one pipeline and delivered dashboards that turn that event stream into behavior insights the platform’s teams can act on.

The payoff

A serverless AWS architecture now processes over a billion user events a day at real-time latency, backed by a 6TB OpenSearch cluster serving 2,000+ queries per second — with the legacy PostgreSQL/MySQL migration behind it and dashboards in front of it for the teams driving personalization and engagement.

Serverless buys elasticity, not accuracy. The real engineering was the dedup and scoring layer underneath — the part that decides whether a real-time recommendation is actually worth showing a user.

Platform & Software Engineering Practice Lead, SquareShift