Diagrammatic

Ride Sharing Service — System Design Interview Practice

Design a ride-sharing platform like Uber or Lyft that matches drivers with passengers in real-time. Work through the requirements, architecture trade-offs, and an interactive design review.

Concepts and architecture decisions to consider

  • location servicesConcept to explore
  • real timeConcept to explore
  • geospatialConcept to explore
  • paymentsConcept to explore
  • mobileConcept to explore

Interview prompt

Design Design a ride-sharing platform like Uber or Lyft that matches drivers with passengers in real-time. so users can Match drivers with passengers reliably at scale.

  • Define the source of truth for Match drivers with passengers; Real-time location tracking and make retries idempotent.
  • Use bounded, partitioned state to meet Handle 10M daily rides and Matching should happen within seconds.
  • Separate the critical request path from Geospatial indexing for location queries, WebSockets for real-time updates, Machine learning for demand prediction.
  • Explain consistency, failure recovery, authorization, observability, and a degraded mode.

Requirements and scale assumptions

  • Support the core workflow to Match drivers with passengers.
  • Expose status, results, and freshness appropriate to Design a ride-sharing platform like Uber or Lyft that matches drivers with passengers in real-time..
  • Support authorization, validation, updates, deletion, and recovery semantics.
  • Meet Matching should happen within seconds under normal load.
  • Scale to Handle 10M daily rides without a single hot key or unbounded synchronous work.
  • Do not lose committed state; make retries and duplicate events safe.
  • Degrade safely when downstream workers, caches, or external dependencies fail.
  • Handle 10M daily rides
  • Partition by the primary tenant, user, item, or geographic key and isolate hot partitions.
  • Keep serving state bounded; retain raw events or durable records for replay and auditing.
  • Peak scale: Handle 10M daily rides — Capacity assumption that drives partitioning and backpressure.
  • Latency target: Matching should happen within seconds — User-facing budget for the primary request or read path.
  • Durable boundary: Committed before async — The source of truth is Match drivers with passengers; Real-time location tracking.
  • Async boundary: At-least-once workers — Keep Geospatial indexing for location queries, WebSockets for real-time updates, Machine learning for demand prediction off the synchronous path.

Key entities

  • InteractioninteractionId, actorId, objectId, type, version, occurredAt

    Canonical ride sharing service interaction with an idempotency key and ordering version.

  • ConnectionSessionsessionId, userId, deviceId, roomKey, lastHeartbeat, status

    Ephemeral but observable ride sharing service connection registration used for routing and presence.

  • FanoutCursorstreamKey, shard, offset, consumerGroup, updatedAt

    Durable progress marker for ride sharing service fan-out and replay.

  • DeliveryReceiptinteractionId, recipientId, channel, attempt, status, deliveredAt

    Deduplicated ride sharing service delivery state for reconnects, retries, or acknowledgements.

Data flow

  1. 1. Accept and commit the interactionThe ride sharing service gateway authenticates the actor, validates room or object membership, applies rate limits, and conditionally commits the interaction.
  2. 2. Publish an ordered eventAn outbox emits the committed ride sharing service transition with an event ID, partition key, sequence, and replay retention.
  3. 3. Fan out by partitionConsumers route ride sharing service events to connected recipients, durable inboxes, or notification channels without making the origin write wait for every recipient.
  4. 4. Resume and reconcile connectionsClients reconnect with a cursor; the ride sharing service service replays missed events, deduplicates delivery, and exposes stale or degraded state.
  5. 5. Measure latency and recoverOperations tracks ride sharing service publish-to-deliver latency, hot partitions, reconnect storms, dropped events, and consumer lag for replay or repair.

Deep dives and trade-offs

  • Ordering, idempotency, and hot keysChoose a ride sharing service partition key that preserves required order while distributing high-volume rooms, users, or objects. Use event IDs, inboxes, consumer offsets, and conditional state transitions for at-least-once delivery. Split or isolate hot partitions without changing the client-visible sequence contract.
  • Reconnect and replay semanticsIssue resumable ride sharing service cursors with an expiry and a clear snapshot-plus-delta fallback. Bound replay windows and rebuild from durable state when a cursor is too old. Expose version and freshness so a client can distinguish current, catching up, and degraded state.
  • Backpressure and presenceKeep connection heartbeats and ephemeral presence separate from durable ride sharing service interactions. Coalesce safe updates, shed low-value work, and protect critical events during reconnect storms. Measure end-to-end delivery, not only broker publish latency.
  • Direct fan-out versus pull-based readsUse push for latency-sensitive ride sharing service deltas and pull or replay for reconnect, history, and recovery. A push-only design loses state when clients disconnect and a pull-only design wastes latency and bandwidth.
  • Per-recipient queues versus shared streamsUse shared partitioned streams with per-recipient cursors where fan-out is large, and isolate exceptional high-fanout objects. A queue per recipient becomes expensive and hard to inspect at large scale.
  • Strong ordering versus availabilityGuarantee ordering only within the scope the product needs, such as a room, object, or conversation. Global ordering introduces a bottleneck and still does not solve duplicate delivery or reconnect recovery.
Diagrammatic — system design practice and architecture review.