Case study
Real-Time Content Discovery on AWS OpenSearch
How SquareShift built a serverless AWS OpenSearch pipeline that powers real-time content discovery and behavioral analytics for a global AI-powered learning platform.
A global AI-powered learning platform used by enterprises and governments, processing more than a billion user events every day.
That real-time behavioral data drives the personalized learning and engagement experience the platform is built around.
Serverless AWS architecture built on Glue and OpenSearch. GB of new content ingested daily. Search queries served per second at peak.
Key servicesHow SquareShift delivered it.
The challenge
The platform needed real-time behavioral analytics with minimal latency, at a scale of over a billion user events a day. Ingestion had to pull from multiple sources at once — including legacy data still sitting in PostgreSQL and MySQL — and deduplicate it without slowing anything down.
What we delivered
SquareShift migrated more than a terabyte of historical data to AWS and built a fully serverless architecture on AWS Glue and OpenSearch. The pipeline now ingests roughly 30GB of new content every day, running through a 6TB OpenSearch cluster tuned to serve over 2,000 queries per second.
SquareShift consolidated the multiple source systems into one pipeline and delivered dashboards that turn that event stream into behavior insights the platform’s teams can act on.
The payoff
A serverless AWS architecture now processes over a billion user events a day at real-time latency, backed by a 6TB OpenSearch cluster serving 2,000+ queries per second — with the legacy PostgreSQL/MySQL migration behind it and dashboards in front of it for the teams driving personalization and engagement.
Serverless buys elasticity, not accuracy. The real engineering was the dedup and scoring layer underneath — the part that decides whether a real-time recommendation is actually worth showing a user.
Platform & Software Engineering Practice Lead, SquareShift
Where this work sits
