<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Hybrid rate limiting]]></title><description><![CDATA[Hybrid rate limiting]]></description><link>https://hybrid-rate-limiting.hashnode.dev</link><generator>RSS for Node</generator><lastBuildDate>Mon, 28 Sep 2026 10:09:23 GMT</lastBuildDate><atom:link href="https://hybrid-rate-limiting.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[⚔️ Hybrid Rate Limiting — The Cloudflare-Style Token + Leaky Bucket Combo]]></title><description><![CDATA[“Token Bucket allows bursts. Leaky Bucket ensures smooth flow.Together — they make your backend bulletproof.”
So far, we’ve understood two classic strategies:

🪣 Token Bucket — Burst-friendly, user-first

💧 Leaky Bucket — Consistent, system-first

...]]></description><link>https://hybrid-rate-limiting.hashnode.dev/hybrid-rate-limiting-the-cloudflare-style-token-leaky-bucket-combo</link><guid isPermaLink="true">https://hybrid-rate-limiting.hashnode.dev/hybrid-rate-limiting-the-cloudflare-style-token-leaky-bucket-combo</guid><category><![CDATA[leakybucket]]></category><category><![CDATA[tokenbucket]]></category><category><![CDATA[backend]]></category><dc:creator><![CDATA[Saurabh Singh Rajput]]></dc:creator><pubDate>Sat, 25 Oct 2025 06:01:39 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1761372077382/518f9ca2-fbd2-46bd-b923-f3f1833c1495.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>“Token Bucket allows bursts. Leaky Bucket ensures smooth flow.<br />Together — they make your backend <em>bulletproof</em>.”</p>
<p>So far, we’ve understood two classic strategies:</p>
<ul>
<li><p>🪣 <strong>Token Bucket</strong> — Burst-friendly, user-first</p>
</li>
<li><p>💧 <strong>Leaky Bucket</strong> — Consistent, system-first</p>
</li>
</ul>
<p>But real-world systems need both.<br />Because real traffic isn’t uniform — it comes in unpredictable <strong>bursts</strong>, <strong>peaks</strong>, and <strong>spam waves</strong>.<br />This is where the <strong>Hybrid Model</strong> shines — combining flexibility + stability.</p>
<hr />
<h2 id="heading-the-real-world-problem">🧠 The Real-World Problem</h2>
<p>Imagine your API is popular.<br />Suddenly, a trending client sends <strong>1000 requests in one second</strong> (due to user activity spike).</p>
<ul>
<li><p>You don’t want to block them immediately (it’s legit traffic).</p>
</li>
<li><p>But you also don’t want your backend to melt down.</p>
</li>
</ul>
<p>You need:</p>
<ul>
<li><p>To <strong>allow short bursts</strong> (good UX)</p>
</li>
<li><p>To <strong>process them at a controlled rate</strong> (backend safety)</p>
</li>
</ul>
<p>That’s what the <strong>Hybrid Rate Limiter</strong> achieves.</p>
<hr />
<h2 id="heading-the-core-idea">🪣 + 💧 The Core Idea</h2>
<blockquote>
<p>“Allow bursts using a Token Bucket —<br />Process them steadily using a Leaky Bucket.”</p>
</blockquote>
<p>In simpler terms:</p>
<ol>
<li><p><strong>Token Bucket (Frontline Layer)</strong></p>
<ul>
<li><p>Decides whether a request is <em>eligible</em> or not</p>
</li>
<li><p>Adds fairness — each user has limited burst tokens</p>
</li>
</ul>
</li>
<li><p><strong>Leaky Bucket (Processing Layer)</strong></p>
<ul>
<li><p>Controls <em>how fast</em> requests actually flow to backend</p>
</li>
<li><p>Ensures steady processing and avoids overload</p>
</li>
</ul>
</li>
</ol>
<p>Together:</p>
<ul>
<li><p>Users can send bursts temporarily</p>
</li>
<li><p>Your backend still processes them at a safe, constant rate</p>
</li>
</ul>
<hr />
<h2 id="heading-architecture-overview">🧩 Architecture Overview</h2>
<pre><code class="lang-javascript">Incoming Requests
       ↓
 [Token Bucket]  →  Filters excessive spam
       ↓
 [Leaky Bucket]  →  Controls processing flow
       ↓
   Application Logic / Database / Cache
</code></pre>
<hr />
<h2 id="heading-algorithm-steps">⚙️ Algorithm Steps</h2>
<ol>
<li><p>A request arrives</p>
</li>
<li><p><strong>Token Bucket check</strong> → if token available → proceed; else reject</p>
</li>
<li><p><strong>Leaky Bucket queue</strong> → add request to steady pipeline</p>
</li>
<li><p><strong>Requests leak (process)</strong> at constant rate downstream</p>
</li>
<li><p>Refill tokens periodically (allowing future bursts)</p>
</li>
</ol>
<p>This pattern is perfect for <strong>distributed systems</strong>, <strong>API gateways</strong>, and <strong>microservices throttling</strong>.</p>
<hr />
<h2 id="heading-implementation-nodejs-redis">💻 Implementation (Node.js + Redis)</h2>
<p>Let’s build this <strong>Cloudflare-style dual layer rate limiter</strong>.</p>
<hr />
<h3 id="heading-install-dependencies">⚙️ Install Dependencies</h3>
<pre><code class="lang-bash">npm install express ioredis
</code></pre>
<hr />
<h3 id="heading-hybridlimiterjs">📁 <code>hybridLimiter.js</code></h3>
<pre><code class="lang-js"><span class="hljs-keyword">import</span> Redis <span class="hljs-keyword">from</span> <span class="hljs-string">"ioredis"</span>;

<span class="hljs-keyword">const</span> redis = <span class="hljs-keyword">new</span> Redis({
  <span class="hljs-attr">host</span>: <span class="hljs-string">"127.0.0.1"</span>,
  <span class="hljs-attr">port</span>: <span class="hljs-number">6379</span>,
});

<span class="hljs-comment">// Token Bucket Parameters</span>
<span class="hljs-keyword">const</span> TOKEN_CAPACITY = <span class="hljs-number">20</span>;    <span class="hljs-comment">// burst capacity</span>
<span class="hljs-keyword">const</span> REFILL_RATE = <span class="hljs-number">1</span>;        <span class="hljs-comment">// 1 token/sec</span>

<span class="hljs-comment">// Leaky Bucket Parameters</span>
<span class="hljs-keyword">const</span> LEAK_CAPACITY = <span class="hljs-number">50</span>;     <span class="hljs-comment">// max queue length</span>
<span class="hljs-keyword">const</span> LEAK_RATE = <span class="hljs-number">5</span>;          <span class="hljs-comment">// steady requests per second</span>

<span class="hljs-keyword">export</span> <span class="hljs-keyword">async</span> <span class="hljs-function"><span class="hljs-keyword">function</span> <span class="hljs-title">hybridRateLimiter</span>(<span class="hljs-params">clientId</span>) </span>{
  <span class="hljs-keyword">const</span> tokenKey = <span class="hljs-string">`token:<span class="hljs-subst">${clientId}</span>`</span>;
  <span class="hljs-keyword">const</span> leakKey = <span class="hljs-string">`leak:<span class="hljs-subst">${clientId}</span>`</span>;
  <span class="hljs-keyword">const</span> now = <span class="hljs-built_in">Date</span>.now();

  <span class="hljs-comment">// --- TOKEN BUCKET CHECK ---</span>
  <span class="hljs-keyword">const</span> tokenBucket = <span class="hljs-keyword">await</span> redis.hgetall(tokenKey);
  <span class="hljs-keyword">let</span> tokens = <span class="hljs-built_in">parseFloat</span>(tokenBucket.tokens || TOKEN_CAPACITY);
  <span class="hljs-keyword">let</span> lastRefill = <span class="hljs-built_in">parseInt</span>(tokenBucket.lastRefill || now);
  <span class="hljs-keyword">const</span> elapsed = (now - lastRefill) / <span class="hljs-number">1000</span>;
  <span class="hljs-keyword">const</span> refill = elapsed * REFILL_RATE;
  tokens = <span class="hljs-built_in">Math</span>.min(TOKEN_CAPACITY, tokens + refill);
  lastRefill = now;

  <span class="hljs-keyword">let</span> tokenAllowed = <span class="hljs-literal">false</span>;
  <span class="hljs-keyword">if</span> (tokens &gt;= <span class="hljs-number">1</span>) {
    tokens -= <span class="hljs-number">1</span>;
    tokenAllowed = <span class="hljs-literal">true</span>;
  }

  <span class="hljs-keyword">await</span> redis.hset(tokenKey, <span class="hljs-string">"tokens"</span>, tokens, <span class="hljs-string">"lastRefill"</span>, lastRefill);
  <span class="hljs-keyword">await</span> redis.expire(tokenKey, <span class="hljs-number">60</span>);

  <span class="hljs-keyword">if</span> (!tokenAllowed) <span class="hljs-keyword">return</span> <span class="hljs-literal">false</span>; <span class="hljs-comment">// reject here itself</span>

  <span class="hljs-comment">// --- LEAKY BUCKET FLOW CONTROL ---</span>
  <span class="hljs-keyword">const</span> leakBucket = <span class="hljs-keyword">await</span> redis.hgetall(leakKey);
  <span class="hljs-keyword">let</span> water = <span class="hljs-built_in">parseFloat</span>(leakBucket.water || <span class="hljs-number">0</span>);
  <span class="hljs-keyword">let</span> lastLeak = <span class="hljs-built_in">parseInt</span>(leakBucket.lastLeak || now);

  <span class="hljs-keyword">const</span> leakElapsed = (now - lastLeak) / <span class="hljs-number">1000</span>;
  <span class="hljs-keyword">const</span> leaked = leakElapsed * LEAK_RATE;
  water = <span class="hljs-built_in">Math</span>.max(<span class="hljs-number">0</span>, water - leaked);
  lastLeak = now;

  <span class="hljs-keyword">let</span> leakAllowed = <span class="hljs-literal">false</span>;
  <span class="hljs-keyword">if</span> (water &lt; LEAK_CAPACITY) {
    water += <span class="hljs-number">1</span>;
    leakAllowed = <span class="hljs-literal">true</span>;
  }

  <span class="hljs-keyword">await</span> redis.hset(leakKey, <span class="hljs-string">"water"</span>, water, <span class="hljs-string">"lastLeak"</span>, lastLeak);
  <span class="hljs-keyword">await</span> redis.expire(leakKey, <span class="hljs-number">60</span>);

  <span class="hljs-keyword">return</span> leakAllowed;
}
</code></pre>
<hr />
<h3 id="heading-appjs">📁 <code>app.js</code></h3>
<pre><code class="lang-js"><span class="hljs-keyword">import</span> express <span class="hljs-keyword">from</span> <span class="hljs-string">"express"</span>;
<span class="hljs-keyword">import</span> { hybridRateLimiter } <span class="hljs-keyword">from</span> <span class="hljs-string">"./hybridLimiter.js"</span>;

<span class="hljs-keyword">const</span> app = express();

<span class="hljs-keyword">async</span> <span class="hljs-function"><span class="hljs-keyword">function</span> <span class="hljs-title">rateLimiter</span>(<span class="hljs-params">req, res, next</span>) </span>{
  <span class="hljs-keyword">const</span> clientId = req.ip; <span class="hljs-comment">// Or use API key for user-based throttling</span>
  <span class="hljs-keyword">const</span> allowed = <span class="hljs-keyword">await</span> hybridRateLimiter(clientId);

  <span class="hljs-keyword">if</span> (!allowed) {
    <span class="hljs-keyword">return</span> res.status(<span class="hljs-number">429</span>).json({
      <span class="hljs-attr">message</span>: <span class="hljs-string">"Too many requests — try again later."</span>,
    });
  }

  next();
}

app.use(rateLimiter);

app.get(<span class="hljs-string">"/api/data"</span>, <span class="hljs-function">(<span class="hljs-params">req, res</span>) =&gt;</span> {
  res.json({ <span class="hljs-attr">message</span>: <span class="hljs-string">"✅ Request successfully processed"</span> });
});

app.listen(<span class="hljs-number">8000</span>, <span class="hljs-function">() =&gt;</span> <span class="hljs-built_in">console</span>.log(<span class="hljs-string">"🚀 Hybrid limiter active on port 8000"</span>));
</code></pre>
<hr />
<h2 id="heading-behavior-explained">🔍 Behavior Explained</h2>
<div class="hn-table">
<table>
<thead>
<tr>
<td>Layer</td><td>Purpose</td><td>Result</td></tr>
</thead>
<tbody>
<tr>
<td><strong>Token Bucket</strong></td><td>Allows short bursts</td><td>Smooth user experience</td></tr>
<tr>
<td><strong>Leaky Bucket</strong></td><td>Controls throughput</td><td>Protects backend</td></tr>
<tr>
<td><strong>Combined Effect</strong></td><td>Balanced, fair, reliable</td><td>Production-grade throttling</td></tr>
</tbody>
</table>
</div><hr />
<h2 id="heading-example-scenario">🧪 Example Scenario</h2>
<div class="hn-table">
<table>
<thead>
<tr>
<td>Time (s)</td><td>Incoming Requests</td><td>Token Bucket</td><td>Leaky Bucket Output</td></tr>
</thead>
<tbody>
<tr>
<td>0</td><td>20</td><td>✅ All allowed</td><td>5 processed</td></tr>
<tr>
<td>1</td><td>10</td><td>✅ Allowed</td><td>5 processed</td></tr>
<tr>
<td>2</td><td>5</td><td>✅ Allowed</td><td>5 processed</td></tr>
<tr>
<td>3</td><td>0</td><td>Tokens refill</td><td>5 processed</td></tr>
<tr>
<td>4</td><td>20</td><td>✅ Allowed (burst absorbed)</td><td>5 processed</td></tr>
</tbody>
</table>
</div><p>👉 Even with bursts, your API remains steady and responsive.</p>
<hr />
<h2 id="heading-real-world-use-cases">⚙️ Real-World Use Cases</h2>
<p>✅ <strong>Cloudflare / NGINX / AWS API Gateway</strong></p>
<ul>
<li><p>Token Bucket: track client credits</p>
</li>
<li><p>Leaky Bucket: forward requests to backend gradually</p>
</li>
</ul>
<p>✅ <strong>Payment / Transaction APIs</strong></p>
<ul>
<li>Handle multiple retries gracefully</li>
</ul>
<p>✅ <strong>AI APIs / Streaming Backends</strong></p>
<ul>
<li>Allow concurrent users but control GPU or model load rate</li>
</ul>
<p>✅ <strong>Microservices with Message Queues</strong></p>
<ul>
<li>Push all accepted requests into Kafka/RabbitMQ → process at stable leak rate</li>
</ul>
<hr />
<h2 id="heading-production-enhancements">🔒 Production Enhancements</h2>
<ol>
<li><p><strong>Dynamic Rate Per Tier</strong></p>
<pre><code class="lang-js"> <span class="hljs-keyword">const</span> TOKEN_CAPACITY = user.plan === <span class="hljs-string">"pro"</span> ? <span class="hljs-number">100</span> : <span class="hljs-number">20</span>;
 <span class="hljs-keyword">const</span> LEAK_RATE = user.plan === <span class="hljs-string">"pro"</span> ? <span class="hljs-number">10</span> : <span class="hljs-number">3</span>;
</code></pre>
</li>
<li><p><strong>Expose Rate Limit Headers</strong></p>
<pre><code class="lang-js"> res.set({
   <span class="hljs-string">"X-RateLimit-Remaining"</span>: <span class="hljs-built_in">Math</span>.floor(tokens),
   <span class="hljs-string">"X-RateLimit-LeakLevel"</span>: <span class="hljs-built_in">Math</span>.floor(water),
 });
</code></pre>
</li>
<li><p><strong>Monitoring</strong></p>
<ul>
<li><p>Use <strong>Prometheus + Grafana</strong> for rate-limited request count</p>
</li>
<li><p>Alert when rejection rate &gt; threshold</p>
</li>
</ul>
</li>
<li><p><strong>Redis Cluster</strong></p>
<ul>
<li>For horizontal scalability across multiple instances</li>
</ul>
</li>
</ol>
<hr />
<h2 id="heading-when-to-use-the-hybrid-approach">🧭 When to Use the Hybrid Approach</h2>
<div class="hn-table">
<table>
<thead>
<tr>
<td>Scenario</td><td>Recommendation</td></tr>
</thead>
<tbody>
<tr>
<td>Bursty traffic + backend safety required</td><td>✅ Hybrid</td></tr>
<tr>
<td>Smooth, predictable traffic</td><td>💧 Leaky only</td></tr>
<tr>
<td>Low-latency APIs with caching</td><td>🪣 Token only</td></tr>
</tbody>
</table>
</div><hr />
<h2 id="heading-tldr-summary">🧠 TL;DR Summary</h2>
<div class="hn-table">
<table>
<thead>
<tr>
<td>Strategy</td><td>Handles Bursts</td><td>Smooth Flow</td><td>Ideal For</td></tr>
</thead>
<tbody>
<tr>
<td>Token Bucket</td><td>✅</td><td>❌</td><td>API clients, burst handling</td></tr>
<tr>
<td>Leaky Bucket</td><td>❌</td><td>✅</td><td>Streaming, message queues</td></tr>
<tr>
<td>Hybrid</td><td>✅</td><td>✅</td><td>Large-scale APIs, Cloudflare-style systems</td></tr>
</tbody>
</table>
</div><hr />
<h2 id="heading-final-thought">💬 Final Thought</h2>
<blockquote>
<p>“Token Bucket is kindness. Leaky Bucket is discipline.<br />Together — they create <em>balance</em>.”</p>
</blockquote>
<p>This hybrid approach is how modern systems gracefully manage <strong>millions of requests per minute</strong> without falling apart — providing reliability, fairness, and control in one elegant design.</p>
<hr />
<h2 id="heading-bonus">🧩 Bonus</h2>
<p>If you’re building a <strong>backend project</strong>, you can:</p>
<ul>
<li><p>Use this hybrid limiter in your Express/FastAPI service</p>
</li>
<li><p>Add Redis-based tracking per user/API key</p>
</li>
<li><p>Visualize rate limit analytics on Grafana</p>
</li>
</ul>
<p>It’s how you transform a simple API into a <strong>production-grade system</strong>. 🚀</p>
<hr />
]]></content:encoded></item></channel></rss>