Back to All Articles
Interview Prep
September 28, 2026
12 min read

System Design Interview Cheat Sheet: Scalability, Trade-Offs, and Architectures

Ace your senior engineering system design loop. Master database sharding, caching strategies, message queues, rate limiters, and the 4-step framework used by Staff Engineers.

H
Himanshu Kumar
Founder & AI Systems Architect, HireOrbitAi

1. The 4-Step System Design Interview Framework#

A 45-minute senior engineering system design interview moves with extreme velocity. Candidate failure is almost never caused by a lack of knowledge regarding databases or caches; it is caused by chaotic time management and wandering without a structured framework.

Staff and Principal Engineers navigate system design evaluations using a disciplined four-step cadence:

Step 1: Clarify Requirements & Scope (5-7 Minutes)

Never start drawing boxes or selecting databases before establishing clear boundaries:

  • Functional Requirements: What 2-3 core features must the architecture support? (e.g., *User can post a 280-character message, follow other accounts, and view a real-time chronological timeline*).
  • Non-Functional Requirements: What are the latency constraints (e.g., read latency < 50ms, write latency < 200ms)? Is high availability prioritized over strict consistency (CAP Theorem)?
  • Traffic & Scale Estimates: Daily Active Users (DAU), read-to-write ratio, peak Queries Per Second (QPS), and 5-year storage projections.
  • Step 2: High-Level Architecture (10-12 Minutes)

    Construct the end-to-end data flow from client devices to backend persistence:

  • Client $\rightarrow$ DNS / CDN $\rightarrow$ Load Balancer (Nginx / ALB) $\rightarrow$ API Gateway $\rightarrow$ Stateless Microservices $\rightarrow$ Primary Database & Caching Layer.
  • Step 3: Deep Dive into Core Bottlenecks (15-20 Minutes)

    The interviewer will probe your design with challenging failure scenarios:

  • *"How does the system handle a viral celebrity posting a tweet to 80 million followers?"*
  • *"What happens when the primary database replica crashes during peak traffic?"*
  • Detail your database sharding keys, indexing strategies, cache invalidation protocols, and fan-out mechanics.
  • Step 4: System Robustness & Failure Modes (5 Minutes)

    Conclude by addressing single points of failure (SPOF), telemetry monitoring, circuit breakers, rate limiting, and graceful degradation strategies.


    2. Essential Back-of-the-Envelope Mental Calculations#

    Memorize these mathematical constants to execute capacity estimates effortlessly on the whiteboard:

  • $1 \text{ day} = 86,400 \text{ seconds} \approx 10^5 \text{ seconds}$
  • $100 \text{ Million DAU} \times 10 \text{ requests/day} = 10^9 \text{ requests/day} \approx 10,000 \text{ QPS}$
  • Peak QPS is typically estimated at $2\times$ to $3\times$ average QPS.
  • Storage Units: $1 \text{ KB} = 10^3 \text{ bytes} \quad | \quad 1 \text{ MB} = 10^6 \text{ bytes} \quad | \quad 1 \text{ GB} = 10^9 \text{ bytes} \quad | \quad 1 \text{ TB} = 10^{12} \text{ bytes}$
  • If each post consumes 500 bytes and users generate 100M posts daily: $10^8 \times 500 = 50 \text{ GB/day} \approx 18.25 \text{ TB/year}$.

  • 3. Storage Engine Trade-Offs: SQL vs. NoSQL vs. NewSQL#

    Scroll table horizontallySwipe ➔
    Architectural DimensionRelational (SQL)Document (NoSQL)Key-Value StoreWide-Column
    Prominent SystemsPostgreSQL, MySQLMongoDB, CouchbaseRedis, MemcachedApache Cassandra, ScyllaDB
    Data SchemaRigid, normalized, ACID compliantFlexible, hierarchical JSONIn-memory key-blobDenormalized, tabular
    Scaling ProfileVertical scaling, read replicas, complex shardingHorizontal sharding out of the boxIn-memory clusteringMassive horizontal write scaling
    Best UtilizationFinancial ledgers, ACID order checkoutCatalogs, content managementSession caching, leaderboardsTime-series telemetry, chat history

    4. Advanced Caching Patterns & Eviction Strategies#

    Caching dramatically reduces database IOPS bottlenecks and slashes p99 latency:

  • Cache-Aside (Lazy Loading): The application queries the cache first. On a cache miss, it reads from the primary database, populates the cache, and returns the result. *Ideal for read-heavy workloads.*
  • Write-Through: The application writes simultaneously to the cache and the database. Ensures data consistency at the expense of slightly higher write latency.
  • Write-Behind (Write-Back): The application writes exclusively to the cache; the cache asynchronously flushes batches to the database. *Provides maximum write throughput, but risks data loss during server crashes.*
  • Cache Invalidation Eviction Strategies: LRU (Least Recently Used), LFU (Least Frequently Used), and strict TTL expiration to prevent stale data.

  • 5. Distributed Coordination: Messaging & Rate Limiting#

  • Message Brokers (Apache Kafka vs. RabbitMQ vs. AWS SQS): Decouple synchronous client HTTP requests from asynchronous heavy processing (email dispatch, video encoding, push notifications).
  • Rate Limiting Algorithms:
  • Token Bucket: Allows bursts up to bucket capacity while refilling at a constant rate. Standard for modern API gateways.
  • Sliding Window Counter: Blends low memory consumption with smooth traffic distribution, preventing burst exploits at window boundaries.

  • 6. Real-World Case Study: Designing a Distributed URL Shortener#

    To demonstrate how the 4-step framework functions during a live interview, consider the classic TinyURL system:

    1. Requirements & Scale

  • Write Volume: 100M URLs generated monthly $\approx 40$ writes/sec.
  • Read Volume: 10:1 read-to-write ratio $\approx 400$ reads/sec.
  • Short URL Length: 7 Base62 characters ($62^7 \approx 3.5$ trillion unique keys).
  • 2. Architectural Choices

  • Database Selection: A distributed Key-Value NoSQL store (DynamoDB or Cassandra) is selected because URL mappings are simple key-value pairs with zero relational joins.
  • Hashing Strategy: Generating Short URLs using a pre-allocated distributed sequence generator (Snowflake ID generator) encoded in Base62 to completely eliminate hash collision retries.
  • Caching Layer: Redis cluster caching the top 20% most active URLs (Pareto Principle), absorbing 80% of read queries and maintaining p99 read latency under 15ms.

  • 7. Latency Numbers Every System Architect Must Know#

    When defending your architectural calculations in Staff-level interviews, anchor your reasoning in fundamental physical hardware constraints:

    Scroll table horizontallySwipe ➔
    OperationApproximate Physical LatencyScaled Human Intuition
    L1 CPU Cache Reference$0.5 \text{ ns}$1 heartbeat
    Branch Mispredict$5 \text{ ns}$10 heartbeats
    L2 CPU Cache Reference$7 \text{ ns}$14 heartbeats
    Main Memory (RAM) Access$100 \text{ ns}$3.3 minutes
    SSD Random Read (NVMe)$150 \ \mu\text{s}$3.5 days
    Datacenter Round-Trip (Same Rack)$500 \ \mu\text{s}$11 days
    Standard HDD Seek$10 \text{ ms}$8 months
    Internet Packet Transatlantic (NYC to London)$150 \text{ ms}$10 years

    Understanding that memory access is nearly a million times faster than disk I/O and network requests is why intelligent caching layers and connection pooling represent the first line of defense in distributed scalability.

    Frequently Asked Questions

    Basic arithmetic using powers of ten and powers of two. Memorize numbers like seconds in a day (86,400 ≈ 100,000), bytes per character, and network latency thresholds.
    Share Article
    HireOrbitAi Power Feature

    Practice mock system design scenarios with AI

    HireOrbitAi's AI Coach challenges your architectural choices, asks edge-case questions, and grades your distributed systems trade-offs.

    Start System Design Prep
    H

    Written by Himanshu Kumar

    Founder & AI Systems Architect, HireOrbitAi

    Building next-generation AI agents and semantic career intelligence platforms. Helping engineers and leaders bridge the gap between technical capability and dream job offers.

    Tags:#System Design#Distributed Systems#Software Architecture#Senior Engineer#Scalability