Top K Elements: App Store Rankings, Amazon Bestsellers — System Design Interview Practice
Design a system to track and display top K items (products, apps, etc.) based on various metrics in real-time. Work through the requirements, architecture trade-offs, and an interactive design review.
Concepts and architecture decisions to consider
- rankingConcept to explore
- top kConcept to explore
- heapConcept to explore
- real timeConcept to explore
- aggregationConcept to explore
Interview prompt
Design Design a system to track and display top K items (products, apps, etc.) based on various metrics in real-time. so users can Track top K items by sales/downloads/ratings reliably at scale.
- Define the source of truth for Track top K items by sales/downloads/ratings; Update rankings in real-time or near real-time and make retries idempotent.
- Use bounded, partitioned state to meet Handle millions of items and Fast ranking updates.
- Separate the critical request path from Min-heap or max-heap for top K, Redis sorted sets for rankings, Periodic batch updates vs real-time.
- Explain consistency, failure recovery, authorization, observability, and a degraded mode.
Requirements and scale assumptions
- Support the core workflow to Track top K items by sales/downloads/ratings.
- Expose status, results, and freshness appropriate to Design a system to track and display top K items (products, apps, etc.) based on various metrics in real-time..
- Support authorization, validation, updates, deletion, and recovery semantics.
- Meet Fast ranking updates under normal load.
- Scale to Handle millions of items without a single hot key or unbounded synchronous work.
- Do not lose committed state; make retries and duplicate events safe.
- Degrade safely when downstream workers, caches, or external dependencies fail.
- Handle millions of items
- Partition by the primary tenant, user, item, or geographic key and isolate hot partitions.
- Keep serving state bounded; retain raw events or durable records for replay and auditing.
- Peak scale: Handle millions of items — Capacity assumption that drives partitioning and backpressure.
- Latency target: Fast ranking updates — User-facing budget for the primary request or read path.
- Durable boundary: Committed before async — The source of truth is Track top K items by sales/downloads/ratings; Update rankings in real-time or near real-time.
- Async boundary: At-least-once workers — Keep Min-heap or max-heap for top K, Redis sorted sets for rankings, Periodic batch updates vs real-time off the synchronous path.
Key entities
- SourcePartitionsourceId, partitionId, cursor, schemaVersion, watermark, status
Replayable top k elements app store rankings amazon bestsellers source evidence and ingestion cursor.
- SchemaVersiondatasetId, version, compatibility, owner, effectiveAt, status
Governed top k elements app store rankings amazon bestsellers contract used to validate producers and consumers.
- ProcessingRunrunId, inputWatermark, checkpoint, qualityStatus, codeVersion, status
Checkpointed top k elements app store rankings amazon bestsellers processing attempt with quality and lineage metadata.
- AnalyticalDatasetdatasetId, partition, watermark, schemaVersion, qualityStatus, location
Curated top k elements app store rankings amazon bestsellers serving partition with freshness and quality state.
Data flow
- 1. Register sources and contractsThe top k elements app store rankings amazon bestsellers catalog records owners, schemas, compatibility rules, retention, lineage, and partitioning before data is accepted.
- 2. Ingest with backpressureConnectors checkpoint top k elements app store rankings amazon bestsellers source cursors, validate schema and deduplication keys, and slow producers when downstream capacity is exhausted.
- 3. Process event time with checkpointsStream or batch engines compute top k elements app store rankings amazon bestsellers transformations using watermarks, late-data policy, state checkpoints, and deterministic code versions.
- 4. Publish quality-gated datasetsOnly top k elements app store rankings amazon bestsellers outputs that pass completeness, freshness, validity, and privacy checks become visible to analytical consumers.
- 5. Serve, replay, and reconcileConsumers read bounded partitions with freshness metadata while operators replay failed top k elements app store rankings amazon bestsellers ranges and compare output checksums.
Deep dives and trade-offs
- Schema evolution and data qualityVersion top k elements app store rankings amazon bestsellers contracts and make compatibility rules explicit for every producer and consumer. Quarantine malformed partitions instead of poisoning the whole dataset. Track row counts, null rates, duplicates, distribution changes, and policy violations by partition.
- Watermarks, late data, and exactly-once effectsUse source cursors and event-time watermarks for top k elements app store rankings amazon bestsellers progress, not wall-clock assumptions. Make checkpoints, output keys, and sink commits retry-safe under at-least-once delivery. Document how late events revise windows, aggregates, or snapshots.
- Replay, lineage, and costKeep immutable top k elements app store rankings amazon bestsellers raw evidence and code or schema versions so failed outputs can be reproduced. Separate hot serving storage from cold retention and cap replay concurrency. Measure freshness, backlog, compute cost, storage growth, and quality-gate failure rate.
- Streaming versus batchUse streaming for freshness-critical top k elements app store rankings amazon bestsellers paths and batch for backfills, compaction, and expensive recomputation. Forcing every workload into streaming makes state, replay, and cost harder to operate.
- Raw retention versus curated-only storageRetain enough immutable raw evidence for replay, audit, and correction, then tier or expire it according to policy. Without raw evidence, a bad transformation can require an unreproducible emergency fix.
- Central warehouse versus domain-owned datasetsCentralize governance and discovery while letting domain owners own contracts and quality signals. A single team owning every transformation becomes a delivery bottleneck and hides data ownership.