Libraries & Ecosystem · کتابخانه‌ها و اکوسیستم متوسطIntermediate ~52 دقیقه مطالعه~46 min read

کش: Caffeine و RedisCaching: Caffeine & Redis

از صفر: چرا کش می‌زنیم، الگوهای cache-aside تا write-behind، سیاست‌های حذف LRU/LFU و W-TinyLFU در Caffeine، تفاوت TTL و TTI، کش محلی (Caffeine) در برابر توزیع‌شده (Redis)، انتزاع کش اسپرینگ، طوفانِ کش و راه‌های مهارش، و معماری دو‌سطحی — با تشبیه، کدِ اجراشدنی و سؤال‌های مصاحبه.From zero: why we cache, the patterns from cache-aside to write-behind, LRU/LFU and Caffeine's W-TinyLFU eviction, TTL vs TTI, local (Caffeine) vs distributed (Redis), Spring's cache abstraction, cache stampedes and how to tame them, and two-tier architecture — with analogies, runnable code, and interview questions.


بیا اول یک حقیقت ساده را بپذیریم: سریع‌ترین کاری که یک برنامه می‌تواند بکند، کاری است که اصلاً انجامش ندهد. اگر جواب یک پرسش را از قبل داشته باشی، دیگر لازم نیست دوباره به دیتابیس بروی، دوباره حساب کنی، یا دوباره از یک سرویس دور بپرسی. کش (cache) دقیقاً همین است: یک دفترچه‌ی کوچک و نزدیک‌به‌دست که جواب‌های پرتکرار را در آن نگه می‌داری تا دفعه‌ی بعد در چند نانوثانیه تحویلشان بدهی، نه در چند میلی‌ثانیه.

در این فصل قرار نیست فقط یاد بگیری روی یک متد @Cacheable بگذاری و رد شوی. قرار است بفهمی کش چه دردی را درمان می‌کند، چرا گاهی خودش می‌شود منبعِ بدترین باگ‌ها، دو ابزار مرجعِ دنیای جاوا — Caffeine برای کشِ محلی و Redis برای کشِ توزیع‌شده — دقیقاً چه‌کار می‌کنند و کِی کدام را انتخاب کنی، و در مصاحبه‌ی سنیور چطور درباره‌ی مبادله‌های سختِ کش حرف بزنی.

نقشه‌ی راه این فصل

مسیری که با هم می‌رویم:

  1. بخش صفر — واژه‌های پایه: hit، miss، eviction، TTL. اول با تشبیه، بعد اصطلاح فنی.
  2. چرا کش — عددهای تأخیر و اینکه چرا کش تفاوتِ «کند» و «سریع» را می‌سازد.
  3. الگوهای کش — cache-aside، read-through، write-through، write-behind، refresh-ahead (با جدول مقایسه).
  4. سیاست‌های حذف — LRU، LFU و چرا Caffeine الگوریتم مدرن W-TinyLFU را انتخاب کرد.
  5. Caffeine در عمل — ساخت کش، LoadingCache، expireAfterWrite در برابر expireAfterAccess، و refreshAfterWrite.
  6. محلی در برابر توزیع‌شده — Caffeine در برابر Redis، و کِی کدام.
  7. Redis و کلاینت‌هایش — Lettuce، Jedis، Redisson و یک نکته‌ی مهمِ لایسنس.
  8. انتزاع کش اسپرینگ@Cacheable، @CacheEvict، @CachePut و پرچمِ نجات‌بخشِ sync.
  9. معماری دو‌سطحی، طوفانِ کش (stampede) و مهارش، بی‌اعتبارسازی، و مبادله‌ی سازگاری.
  10. دام‌ها، بهترین‌شیوه‌ها و سؤال‌های مصاحبه با پاسخ کامل.

بخش صفر — چند واژه که پیش از هر کد باید حسشان کنی

قبل از اینکه سراغ کد برویم، چند اصطلاح هست که در کل فصل برمی‌گردند. بگذار همین اول با تشبیه در ذهنت جا بیفتند.

کش مثل میزِ کارِ یک آشپز

یک آشپزِ حرفه‌ای همه‌ی مواد را از انبارِ سردخانه (دیتابیس) برنمی‌دارد؛ پرمصرف‌ترین‌ها — نمک، روغن، پیاز خردشده — را روی میزِ کنارِ دستش (کش) نگه می‌دارد. رفتن به سردخانه چند ثانیه طول می‌کشد؛ دست‌بردن به میز، یک لحظه. اما میز جا ندارد که کلِ انبار را رویش بریزی، پس آشپز مدام تصمیم می‌گیرد چه چیزی روی میز بماند و چه چیزی برگردد سردخانه. کل داستانِ کش همین است: یک فضای کوچکِ گران‌قیمتِ سریع، و هنرِ تصمیم‌گیری درباره‌ی اینکه چه چیزی لایقِ ماندن روی آن است.

چند واژه که از این‌جا به بعد بی‌وقفه می‌بینی:

  • hit (اصابت): یعنی جوابی که می‌خواستی، در کش بود و مستقیم تحویلت داد. مثل اینکه نمک را روی میز پیدا کنی.
  • miss (فقدان): یعنی جواب در کش نبود، پس مجبور شدی بروی سراغ منبعِ اصلی (دیتابیس، سرویس، محاسبه). مثل رفتن به سردخانه.
  • hit ratio (نرخ اصابت): درصدِ درخواست‌هایی که hit شدند. کشی که ۹۵٪ hit ratio دارد یعنی از هر ۱۰۰ درخواست، ۹۵ تا اصلاً به دیتابیس نرسیدند. این مهم‌ترین عددِ سلامتِ یک کش است.
  • eviction (حذف/بیرون‌اندازی): وقتی کش پر می‌شود، باید چیزی را بیرون بیندازد تا جا برای تازه‌وارد باز شود. کدام را بیرون بیندازد، همان «سیاست حذف» است که جانِ این فصل است.
  • TTL — Time To Live (زمانِ زندگی): یعنی یک ورودی حداکثر چند مدت اجازه دارد در کش بماند، فارغ از اینکه چند بار خوانده شود. مثل تاریخِ انقضای روی بسته‌ی شیر.
  • evict در برابر expire: «expire» یعنی زمانش تمام شد (TTL)، «evict» یعنی جا کم آمد و بیرونش انداختیم. دو دلیلِ متفاوت برای رفتن.
  • stale (بیات): داده‌ای که در کش هست اما دیگر با منبعِ اصلی هم‌خوان نیست — مثل قیمتی که در کش ۱۰۰ است ولی در دیتابیس شده ۱۲۰. مبادله‌ی اصلیِ کش، همیشه با همین «بیات‌بودن» است.
دو مسئله‌ی سختِ علوم کامپیوتر

یک شوخیِ معروف در مهندسی نرم‌افزار می‌گوید: «در علوم کامپیوتر فقط دو مسئله‌ی سخت وجود دارد: بی‌اعتبارسازیِ کش (cache invalidation) و نام‌گذاریِ چیزها.» این شوخی است اما ته‌اش جدی است: افزودنِ کش آسان است؛ مطمئن‌شدن از اینکه کش هیچ‌وقت داده‌ی غلط تحویل ندهد، یکی از سخت‌ترین کارهای مهندسی است. کلِ نیمه‌ی دومِ این فصل درباره‌ی همین سختی است.


چرا اصلاً کش می‌زنیم؟ داستانِ اعداد

کش یک بهینه‌سازیِ تزئینی نیست؛ از دلِ یک واقعیتِ فیزیکی می‌آید: همه‌ی حافظه‌ها هم‌سرعت نیستند. بین خواندن از حافظه‌ی محلیِ برنامه و رفتن به یک دیتابیسِ روی شبکه، چند مرتبه‌ی بزرگیِ (order of magnitude) تفاوت وجود دارد.

عملیات تأخیرِ تقریبی تشبیه در مقیاس انسانی
خواندن از کش محلی درون‌فرایندی (Caffeine) ~۱۰۰ نانوثانیه برداشتن چیزی از روی میزت
رفت‌وبرگشت به Redis روی همان شبکه ~۰٫۵ تا ۱ میلی‌ثانیه پرسیدن از همکارِ اتاقِ بغل
کوئریِ ساده به دیتابیسِ SQL ~۵ تا ۳۰ میلی‌ثانیه رفتن به بایگانیِ طبقه‌ی پایین
فراخوانیِ یک سرویسِ خارجی روی اینترنت ~۵۰ تا ۵۰۰ میلی‌ثانیه نامه‌نگاری با یک شهرِ دیگر

اختلافِ کشِ محلی و دیتابیس حدودِ پنجاه‌هزار برابر است. حالا تصور کن یک صفحه‌ی محصول در هر ثانیه هزار بار باز می‌شود و هر بار همان اطلاعاتِ ثابتِ محصول را از دیتابیس می‌خوانَد. اگر آن را کش کنی، هزار کوئری در ثانیه به تقریباً صفر می‌رسد و دیتابیس نفس می‌کشد.

کش دو چیز را هم‌زمان نجات می‌دهد

اول تأخیر (latency) را: کاربر جواب را سریع‌تر می‌گیرد. دوم بار (load) را: منبعِ اصلی (که معمولاً گران‌ترین و شکننده‌ترین بخشِ سیستم است، مثل دیتابیس) از زیرِ فشارِ درخواست‌های تکراری بیرون می‌آید. خیلی وقت‌ها هدفِ اصلیِ کش نه سرعتِ کاربر، بلکه محافظت از دیتابیس در برابر فروپاشی است. این تمایز را در مصاحبه یادت باشد.

اما همان‌طور که گفتیم، این سرعت رایگان نیست. لحظه‌ای که یک کپی از داده را جای دیگری نگه می‌داری، دو نسخه از حقیقت داری — و هر جا دو نسخه از حقیقت باشد، احتمالِ ناهم‌خوانی هست. پس بیا اول ببینیم چطور داده وارد و خارجِ کش می‌شود.


الگوهای کش: داده چطور جابه‌جا می‌شود

اینکه چه کسی مسئولِ پرکردنِ کش و نوشتن در دیتابیس است، الگوهای مختلفی می‌سازد. این‌ها را با هم اشتباه‌گرفتن، منشأِ خیلی از سردرگمی‌هاست.

۱) Cache-Aside (کنارگذر / lazy loading)

رایج‌ترین الگو، و همانی که وقتی دستی کش می‌زنی معمولاً می‌نویسی. اینجا کدِ برنامه مسئولِ همه‌چیز است؛ کش فقط یک انبارِ ساده‌ی کلید-مقدار است که کنارِ دستت ایستاده (به همین خاطر «کنارگذر»).

cache-aside مثل یخچالِ خانه

وقتی گرسنه‌ای اول درِ یخچال را باز می‌کنی. اگر غذا بود (hit)، می‌خوری. اگر نبود (miss)، خودت می‌روی سوپرمارکت، می‌خری، و یک نسخه هم می‌گذاری در یخچال برای دفعه‌ی بعد. یخچال هیچ‌وقت خودش نمی‌رود خرید؛ تو راوی و کارگردانِ ماجرایی. این دقیقاً cache-aside است.

public Product getProduct(long id) {
    // ۱) اول کش را نگاه کن
    Product cached = cache.getIfPresent(id);
    if (cached != null) {
        return cached;              // hit
    }
    // ۲) miss — برو منبع اصلی
    Product fromDb = productRepository.findById(id);
    // ۳) نتیجه را برای دفعه‌ی بعد در کش بگذار
    if (fromDb != null) {
        cache.put(id, fromDb);
    }
    return fromDb;
}

نقطه‌قوتش سادگی و کنترلِ کامل است؛ نقطه‌ضعفش این است که این «الگوی سه‌مرحله‌ای» را باید همه‌جا دستی تکرار کنی و اگر یادت برود، باگ می‌خوری. مشکلِ بزرگ‌ترش که بعداً می‌بینیم: اگر هزار درخواست هم‌زمان miss بخورند، هر هزار تا با هم به دیتابیس هجوم می‌برند (همان طوفانِ کش).

۲) Read-Through (خواندنِ ازطریق)

اینجا دیگر خودت آن سه مرحله را نمی‌نویسی؛ به کش می‌گویی «اگر نداشتی، این تابع را صدا بزن و خودت پرش کن». کش واسطه‌ی خواندن می‌شود. LoadingCache در Caffeine دقیقاً همین است.

LoadingCache<Long, Product> cache = Caffeine.newBuilder()
    .maximumSize(10_000)
    .build(id -> productRepository.findById(id));  // «تابعِ لود»

// حالا فقط می‌گویی:
Product p = cache.get(42L);   // اگر نبود، خودِ کش تابعِ لود را صدا می‌زند و می‌گذارد

تفاوتِ ظریف اما مهم با cache-aside: منطقِ «چطور پر شود» یک‌جا کنارِ کش زندگی می‌کند، نه پخش‌شده در همه‌ی فراخوانی‌ها. و مهم‌تر — همان‌طور که خواهیم دید — کشِ لودشونده معمولاً درخواست‌های هم‌زمان روی یک کلید را در هم ادغام (coalesce) می‌کند، یعنی فقط یک بار تابعِ لود اجرا می‌شود حتی اگر هزار thread هم‌زمان همان کلید را بخواهند.

۳) Write-Through (نوشتنِ ازطریق)

قرینه‌ی read-through برای نوشتن: هر بار می‌نویسی، هم‌زمان و به‌صورت هم‌زمان (synchronous) هم کش و هم دیتابیس را به‌روز می‌کنی. مزیت: کش هیچ‌وقت از دیتابیس عقب نمی‌افتد. هزینه: هر نوشتن باید منتظرِ دیتابیس بماند، پس نوشتن کند می‌شود.

۴) Write-Behind / Write-Back (نوشتنِ باتأخیر)

اینجا اول در کش می‌نویسی و فوراً برمی‌گردی؛ نوشتن در دیتابیس بعداً، به‌صورت غیرهم‌زمان و اغلب دسته‌ای (batch)، انجام می‌شود.

write-behind مثل صندوق‌دارِ فروشگاه

صندوق‌دار پولت را می‌گیرد و رسید می‌دهد و تو می‌روی (نوشتن در کش، سریع). خودِ پول‌ها را در طول روز جمع می‌کند و آخرِ شب یک‌جا به بانک می‌برد (نوشتن دسته‌ای در دیتابیس). خیلی سریع است، اما یک ریسک دارد: اگر بین‌راه فروشگاه آتش بگیرد (کرشِ سرور)، پول‌هایی که هنوز به بانک نرفته‌اند از بین می‌روند. write-behind هم همین است: سریع، اما در صورت کرش خطرِ ازدست‌رفتنِ نوشته‌های در صف را دارد.

۵) Refresh-Ahead (تازه‌سازیِ پیش‌دستانه)

الگویی که با اسم گول‌نزنی: کش پیش از آنکه یک ورودیِ «داغ» منقضی شود، در پس‌زمینه دوباره لودش می‌کند تا کاربر هیچ‌وقت به miss نخورد. Caffeine این را با refreshAfterWrite می‌دهد که جلوتر می‌بینیم.

دو مبادله‌ای که این الگوها را از هم جدا می‌کند

دو محورِ اصلی را در ذهن بسپار: تازگیِ داده در برابر سرعتِ نوشتن، و سادگیِ کد در برابر مهارِ خودکار. write-through تازه‌ترین است ولی نوشتن را کند می‌کند؛ write-behind سریع‌ترین نوشتن را دارد ولی خطرِ ازدست‌رفتنِ داده و پیچیدگی؛ cache-aside ساده‌ترین ذهنیت را دارد ولی هیچ محافظتِ خودکاری در برابر طوفان ندارد؛ read-through/refresh-ahead این محافظت را می‌دهند به قیمتِ اینکه منطق را به کش بسپاری. انتخابِ درست به این بستگی دارد که کدام‌شان برایت گران‌تر است.

الگو چه کسی می‌نویسد/می‌خوانَد تازگی ریسکِ اصلی کِی مناسب است
Cache-aside کدِ برنامه، دستی متوسط طوفانِ کش، فراموشیِ به‌روزرسانی کارِ عمومی، کنترلِ کامل بخواهی
Read-through کش، با تابعِ لود متوسط وابستگی به کتابخانه خواندن‌محور، ادغامِ خودکار بخواهی
Write-through کش، هم‌زمان بالا (همیشه تازه) نوشتنِ کند جایی که بیاتی غیرقابل‌قبول است
Write-behind کش، باتأخیر پایین (موقتاً) ازدست‌رفتنِ داده در کرش نوشتنِ پرحجم، تحملِ ازدست‌رفتنِ کم
Refresh-ahead کش، پیش‌دستانه در پس‌زمینه بالا برای داغ‌ها لودِ اضافه برای داده‌ی بی‌مصرف کلیدهای داغِ گران‌قیمت

سیاست‌های حذف: وقتی کش پر می‌شود، کدام می‌رود؟

کش محدود است. وقتی پر شد و ورودیِ تازه‌ای آمد، باید یکی را قربانی کند. کدام را قربانی کنی، مستقیماً hit ratio را تعیین می‌کند — و همین است که Caffeine را از یک HashMap ساده جدا می‌کند.

LRU — کم‌اخیراً‌استفاده‌شده

LRU مثل تمیزکردنِ کمدِ لباس

سیاستِ ساده و پرطرفدارِ LRU (Least Recently Used) می‌گوید: «هرچه مدتِ بیشتری است بهش دست نزده‌ای، اول برود.» مثل کمدِ لباست: لباسی که یک سال است نپوشیده‌ای، اولین گزینه‌ی دورانداختن است. منطق ساده و اغلب خوب است: چیزی که تازگی استفاده شده، احتمالاً زود دوباره استفاده می‌شود (اصلِ «محلی‌بودنِ زمانی»).

اما LRU یک ضعفِ مشهور دارد: در برابرِ پویشِ ترتیبی (scan) فرو می‌پاشد. تصور کن کشِ پرِ داده‌های داغت را داری، و یک‌بار کسی کلِ جدول را یک‌دور می‌خواند (مثلاً یک گزارشِ سنگین). این پویش، هزاران ورودیِ یک‌بارمصرف را وارد می‌کند که هر کدام «تازه استفاده‌شده» حساب می‌شوند و همه‌ی داده‌های داغِ واقعی‌ات را از کش بیرون می‌ریزند. LRU نمی‌فهمد که آن‌ها فقط یک بار خواسته شدند.

LFU — کم‌بسامد‌استفاده‌شده

LFU (Least Frequently Used) می‌گوید: «هرچه کم‌تر استفاده شده، اول برود» — یعنی بر اساسِ تعدادِ دفعات تصمیم می‌گیرد، نه زمانِ آخرین دفعه. این در برابرِ پویش مقاوم‌تر است (چون داده‌ی یک‌بارمصرف بسامدِ ۱ دارد و زود قربانی می‌شود). اما LFU هم دو ضعف دارد: نگه‌داشتنِ شمارنده برای همه‌چیز حافظه می‌خواهد، و به گذشته گیر می‌کند — چیزی که ماهِ پیش خیلی داغ بود ولی حالا سرد شده، هنوز شمارنده‌ی بالایی دارد و لجوجانه در کش می‌ماند.

W-TinyLFU — انتخابِ هوشمندانه‌ی Caffeine

اینجاست که Caffeine درخشش می‌کند. Caffeine — که در اصل بازنویسیِ کشِ Guava است و امروز کتابخانه‌ی مرجعِ کشِ محلی در جاواست (نسخه‌ی پایدارِ ۳.x، نیازمندِ جاوا ۱۱ به بالا) — از الگوریتمی به نام W-TinyLFU استفاده می‌کند که بهترینِ هر دو دنیا را می‌گیرد.

ایده‌ی مرکزیِ W-TinyLFU

مشکل با LFU خالص این بود که برای شمردنِ بسامدِ همه‌چیز حافظه‌ی زیادی لازم است. TinyLFU این را حل می‌کند: به‌جای شمارنده‌ی دقیق، از یک ساختارِ فشرده‌ی احتمالاتی به نام frequency sketch (نوعی Count-Min Sketch) استفاده می‌کند که با چند بیت به‌ازای هر کلید، تخمینِ بسامد را نگه می‌دارد. «W» هم یعنی Window: یک پنجره‌ی کوچکِ ورودی که تازه‌واردها اول به آن می‌روند تا شانسِ نشان‌دادنِ خودشان را داشته باشند.

جریانِ کارِ W-TinyLFU این‌طور است — و ارزشِ فهمیدن دارد چون در مصاحبه‌ی سنیور طلاست:

  1. پنجره‌ی ورودی (window): هر ورودیِ تازه اول وارد یک LRUِ کوچک (حدود ۱٪ کش) می‌شود. این به داده‌های تازه فرصت می‌دهد بدونِ اینکه فوراً با کهنه‌کارها رقابت کنند، خودی نشان دهند.
  2. دربانِ پذیرش (admission): وقتی ورودی می‌خواهد از پنجره به فضای اصلی برود، TinyLFU مثل یک دربان قضاوت می‌کند: «آیا بسامدِ تخمینیِ این تازه‌وارد، از بسامدِ قربانیِ فعلیِ فضای اصلی بیشتر است؟» اگر بله، وارد می‌شود و آن قربانی می‌رود؛ اگر نه، خودِ تازه‌وارد رد می‌شود. این همان چیزی است که پویش را خنثی می‌کند — داده‌ی یک‌بارمصرف بسامدِ پایینی دارد و دربان راهش نمی‌دهد.
  3. فضای اصلیِ SLRU: فضای اصلی خودش یک Segmented LRU است با دو بخش: یک بخشِ «آزمایشی (probation)» و یک بخشِ «محافظت‌شده (protected)». ورودی‌ای که دوباره خوانده شود از آزمایشی به محافظت‌شده ارتقا می‌یابد. این ترکیبِ بسامد (چقدر) و اخیربودن (کِی) را با هم درنظر می‌گیرد.
  4. پیرشدنِ شمارنده‌ها: frequency sketch به‌مرورِ زمان همه‌ی شمارنده‌ها را نصف می‌کند (aging)، تا آن مشکلِ LFU که «قهرمانِ ماهِ پیش لجوجانه می‌مانَد» حل شود. داغیِ گذشته کم‌کم فراموش می‌شود.

نتیجه: Caffeine در بارهای واقعی نرخِ اصابتی نزدیک به الگوریتمِ ایده‌آلِ آینده‌بین (Bélády) می‌گیرد، در حالی که سربارِ حافظه‌اش ناچیز است. تو معمولاً هیچ‌کدامِ این‌ها را دستی تنظیم نمی‌کنی؛ فقط یک maximumSize می‌دهی و Caffeine بقیه را هوشمندانه اداره می‌کند.

چرا نباید خودت LRU بنویسی

وسوسه می‌شوی با LinkedHashMap یک LRUِ دست‌ساز بسازی. برای اسباب‌بازی اشکالی ندارد، اما در تولید سه مشکل داری: hit ratioِ ضعیف‌تر در برابرِ پویش، قفل‌گذاریِ سراسری که همزمانی را می‌کُشد، و نبودِ TTL و آمار. Caffeine هر سه را حل کرده و شدیداً بهینه‌شده است — بازنویسیِ آن، تقریباً همیشه اشتباه است.


Caffeine در عمل: ساخت و تنظیم

بیا از یک کشِ ساده شروع کنیم و لایه‌لایه غنی‌اش کنیم. وابستگی (نسخه‌ی پایدارِ فعلی از خانواده‌ی ۳.x):

<dependency>
    <groupId>com.github.ben-manes.caffeine</groupId>
    <artifactId>caffeine</artifactId>
    <version>3.2.2</version>
</dependency>

یک کشِ دستی (cache-aside)، با سقفِ اندازه و آمار:

import com.github.benmanes.caffeine.cache.Cache;
import com.github.benmanes.caffeine.cache.Caffeine;
import java.time.Duration;

Cache<Long, Product> cache = Caffeine.newBuilder()
    .maximumSize(10_000)                     // حذفِ مبتنی‌بر اندازه با W-TinyLFU
    .expireAfterWrite(Duration.ofMinutes(5)) // TTL
    .recordStats()                           // شمارش hit/miss برای پایش
    .build();

Product p = cache.getIfPresent(42L);         // ممکن است null باشد (miss)
cache.put(42L, product);

LoadingCache — نسخه‌ی read-through

اگر تابعِ لود را به Caffeine بدهی، دیگر خودت آن سه مرحله‌ی cache-aside را نمی‌نویسی و — مهم‌تر — ادغامِ خودکارِ درخواست‌ها را رایگان می‌گیری:

import com.github.benmanes.caffeine.cache.LoadingCache;

LoadingCache<Long, Product> cache = Caffeine.newBuilder()
    .maximumSize(10_000)
    .expireAfterWrite(Duration.ofMinutes(5))
    .build(id -> productRepository.findById(id));   // تابعِ لود

Product p = cache.get(42L);   // اگر نبود، لود می‌کند و می‌گذارد؛ هرگز null نیست مگر لود null بدهد
چرا LoadingCache در برابرِ طوفانِ کش امن‌تر است

فرض کن کلیدِ ۴۲ در کش نیست و هم‌زمان هزار thread صدایش می‌زنند. با cache-aindeی دستی، هر هزار تا miss می‌بینند و هر هزار تا به دیتابیس می‌روند. اما LoadingCache تضمین می‌کند برای یک کلیدِ مشخص، فقط یک thread تابعِ لود را اجرا کند؛ بقیه پشتِ همان محاسبه صف می‌کشند و نتیجه‌ی آماده را می‌گیرند. به این «ادغام درخواست (request coalescing)» یا «single-flight» می‌گویند و اولین سپرِ تو در برابرِ طوفان است — البته فقط درونِ همین یک JVM.

TTL در برابر TTI: expireAfterWrite در برابر expireAfterAccess

اینجا یکی از پرتکرارترین سوءتفاهم‌های مصاحبه است. دو نوع انقضای زمانی داریم:

TTL مثل شیر، TTI مثل اشتراکِ باشگاه

TTL (که Caffeine آن را expireAfterWrite می‌نامد) مثل تاریخِ انقضای روی بسته‌ی شیر است: از لحظه‌ای که نوشته شد، ساعت شروع به تیک‌تیک می‌کند و فارغ از اینکه چند بار درش را باز کنی، سرِ موعد فاسد می‌شود. TTI یعنی Time To Idle (که Caffeine آن را expireAfterAccess می‌نامد) مثل اشتراکِ باشگاه است که «اگر ۳۰ روز نیایی باطل می‌شود»: هر بار که استفاده کنی، ساعت ریست می‌شود؛ فقط بی‌مصرف‌ماندنِ طولانی می‌کُشدش.

Cache<String, Session> sessions = Caffeine.newBuilder()
    .expireAfterAccess(Duration.ofMinutes(30))   // TTI: تا وقتی فعال است زنده بماند
    .build();

Cache<String, ExchangeRate> rates = Caffeine.newBuilder()
    .expireAfterWrite(Duration.ofMinutes(1))     // TTL: نرخ ارز هر دقیقه بی‌ارزش می‌شود
    .build();

قاعده‌ی سرانگشتی: برای داده‌ای که در منبع کهنه می‌شود (نرخ ارز، قیمت، موجودی) از expireAfterWrite/TTL استفاده کن — مهم نیست چند بار خواندیش، بعد از یک دقیقه دیگر قابل‌اعتماد نیست. برای داده‌ای که فقط وقتی رهاشده باید پاک شود (نشستِ کاربر، دادهٔ موقتی) از expireAfterAccess/TTI استفاده کن.

refreshAfterWrite: تفاوتِ ظریف با انقضا

refreshAfterWrite با expireAfterWrite فرق دارد — این تله‌ی مصاحبه است

expireAfterWrite سخت‌گیر است: بعد از موعد، ورودی حذف می‌شود و درخواستِ بعدی باید منتظرِ لودِ تازه بماند (یک miss با تأخیر). اما refreshAfterWrite نرم است: بعد از موعد، ورودیِ کهنه هنوز آن‌جاست و فوراً به درخواست‌کننده داده می‌شود، اما در همان لحظه یک لودِ تازه در پس‌زمینه راه می‌افتد تا دفعه‌ی بعد تازه باشد. یعنی refresh کاربر را منتظر نمی‌گذارد و می‌تواند یک بیاتیِ لحظه‌ای بدهد؛ expire تازگی را تضمین می‌کند اما با هزینه‌ی تأخیر. اغلب هر دو را با هم می‌گذاری: refresh کوتاه‌تر از expire.

LoadingCache<Long, Product> cache = Caffeine.newBuilder()
    .refreshAfterWrite(Duration.ofMinutes(1))    // بعد از ۱ دقیقه، در پس‌زمینه تازه کن
    .expireAfterWrite(Duration.ofMinutes(10))    // ولی بعد از ۱۰ دقیقه دیگر بیات را نده
    .build(id -> productRepository.findById(id));

AsyncLoadingCache و کشِ ناهم‌زمان

اگر در دنیای ری‌اکتیو یا CompletableFuture کار می‌کنی، Caffeine نسخه‌ی ناهم‌زمان هم دارد که نتیجه را به‌صورتِ CompletableFuture نگه می‌دارد — و همین باعث می‌شود ادغامِ درخواست به‌شکلِ طبیعی روی futureها کار کند:

AsyncLoadingCache<Long, Product> cache = Caffeine.newBuilder()
    .maximumSize(10_000)
    .buildAsync((id, executor) ->
        CompletableFuture.supplyAsync(() -> productRepository.findById(id), executor));

CompletableFuture<Product> future = cache.get(42L);
همیشه recordStats و maximumSize را جدی بگیر

دو تنظیمی که مهندس‌های تازه‌کار فراموش می‌کنند: بدونِ maximumSize یا maximumWeight، کشت یک نشتِ حافظه‌ی بالقوه است که تا OutOfMemoryError رشد می‌کند. و بدونِ recordStats() کور هستی — نمی‌دانی hit ratioات چند است، پس نمی‌دانی کشت اصلاً فایده دارد یا فقط حافظه هدر می‌دهد. کشِ بدونِ پایش، حدس‌زدن است نه مهندسی.


محلی در برابر توزیع‌شده: Caffeine در برابر Redis

تا اینجا از Caffeine حرف زدیم که یک کشِ محلیِ درون‌فرایندی (in-process) است: داده درست درونِ حافظه‌ی همان JVM زندگی می‌کند. سریع‌ترین چیزِ ممکن. اما یک مشکلِ بزرگ دارد که فقط وقتی چند نمونه (instance) از برنامه‌ات را بالا می‌آوری خودش را نشان می‌دهد.

کش محلی مثل دفترچه‌ی جیبیِ هر کارمند

تصور کن یک شرکت ده کارمند دارد و هرکدام یک دفترچه‌ی جیبی (کش محلی) برای یادداشتِ قیمت‌ها. سریع است — هر کس فوری به دفترچه‌ی خودش نگاه می‌کند. اما وقتی قیمت عوض می‌شود، دفترچه‌ی کارمندِ اول به‌روز می‌شود و نُه دفترچه‌ی دیگر هنوز قیمتِ قدیم را دارند! حالا مشتری‌ها بسته به اینکه با کدام کارمند حرف بزنند، جوابِ متفاوت می‌گیرند. این «ناهم‌خوانی بین نمونه‌ها» بزرگ‌ترین ضعفِ کش محلی است. راهِ حل: یک دفترِ مرکزیِ مشترک که همه به آن نگاه کنند — این همان Redis است.

Redis یک انبارِ داده‌ی کلید-مقدارِ درون‌حافظه‌ایِ بیرونی و مشترک است که روی شبکه زندگی می‌کند. همه‌ی نمونه‌های برنامه‌ات به یک Redisِ واحد وصل می‌شوند، پس یک نسخه‌ی واحد از حقیقت دارند. هزینه‌اش: هر دسترسی حالا یک رفت‌وبرگشتِ شبکه‌ای است (میلی‌ثانیه، نه نانوثانیه) و سریال‌سازیِ داده (تبدیل آبجکت جاوا به بایت و برعکس).

ویژگی کش محلی (Caffeine) کش توزیع‌شده (Redis)
محلِ داده درونِ حافظه‌ی همان JVM سرورِ جداگانه روی شبکه
سرعتِ دسترسی ~۱۰۰ نانوثانیه ~۰٫۵ تا ۱ میلی‌ثانیه
اشتراک بین نمونه‌ها ندارد؛ هر نمونه جدا دارد؛ یک منبعِ مشترک
سریال‌سازی لازم است؟ نه (آبجکت خام) بله (بایت روی سیم)
با کرشِ برنامه چه می‌شود؟ کش هم می‌رود کش می‌مانَد (بیرونی است)
ظرفیت محدود به heapِ برنامه بسیار بزرگ‌تر، مستقل
قابلیت‌های اضافه فقط کلید-مقدار ساختارهای داده، pub/sub، TTL سمتِ سرور، اسکریپت
قاعده‌ی طلاییِ انتخاب

اگر داده فقط-خواندنیِ پرتکرار و تحملِ ناهم‌خوانیِ کوتاه‌مدت داری، Caffeine بگذار — سریع‌ترین است. اگر چند نمونه داری و به یک نسخه‌ی مشترک و هماهنگ نیاز داری، یا داده باید از کرشِ برنامه جان به‌در ببرد، Redis بگذار. و در سیستم‌های جدی، اغلب هر دو را با هم می‌گذاری: Caffeine به‌عنوان لایه‌ی اول (L1) و Redis به‌عنوان لایه‌ی دوم (L2) — که کمی جلوتر معماری‌اش را می‌بینیم.


Redis و کلاینت‌های جاوایش

Redis (مخففِ REmote DIctionary Server) فقط یک کشِ کلید-مقدار نیست؛ یک انبارِ داده‌ی درون‌حافظه‌ای با ساختارهای داده‌ی غنی است: رشته، هش، لیست، مجموعه، sorted set، و بیشتر. اما پرکاربردترین نقشش هنوز همان کش است.

تک‌رشته‌ای‌بودنِ Redis یک ویژگی است، نه ضعف

شاید بشنوی «Redis تک‌رشته‌ای است». منظور این است که اجرای فرمان‌ها روی یک رشته‌ی منطقیِ واحد سریال می‌شود (نسخه‌های جدید I/O شبکه را چندرشته‌ای کرده‌اند، اما پردازشِ فرمان همچنان سریال است). این عمدی است: چون فقط یک فرمان در آنِ واحد اجرا می‌شود، عملیاتِ Redis به‌صورتِ طبیعی اتمیک هستند و نیازی به قفلِ پیچیده نیست. برای همین دستوری مثل INCR بی‌هیچ race condition کار می‌کند. سریع‌بودنش هم از همین سادگی و درون‌حافظه‌ای‌بودن می‌آید.

سه کلاینتِ اصلیِ جاوا که باید بشناسی:

  • Lettuce: کلاینتِ پیش‌فرضِ اسپرینگ‌بوت. مبتنی بر Netty، ناهم‌زمان و ری‌اکتیو، thread-safe؛ یک اتصال را می‌شود بینِ چند thread به‌اشتراک گذاشت. انتخابِ پیش‌فرضِ خوب.
  • Jedis: کلاینتِ قدیمی‌تر و ساده‌تر، مسدودکننده (blocking) و همگام. thread-safe نیست، پس معمولاً با یک استخرِ اتصال (connection pool) استفاده می‌شود. ساده و شناخته‌شده.
  • Redisson: سطحِ بالاتر؛ به‌جای فرمان‌های خام، پیاده‌سازی‌های توزیع‌شده‌ی ساختارهای جاوا (RMap، RLock، سمافور توزیع‌شده، صف) می‌دهد. وقتی به قفلِ توزیع‌شده یا ساختارهای پیچیده نیاز داری عالی است.

استفاده‌ی ساده با Lettuce (مستقیم، بدونِ اسپرینگ):

import io.lettuce.core.RedisClient;
import io.lettuce.core.api.StatefulRedisConnection;
import io.lettuce.core.api.sync.RedisCommands;

RedisClient client = RedisClient.create("redis://localhost:6379");
try (StatefulRedisConnection<String, String> conn = client.connect()) {
    RedisCommands<String, String> cmd = conn.sync();

    cmd.set("product:42", jsonPayload);      // نوشتن
    cmd.expire("product:42", 300);           // TTL سمتِ سرور: ۳۰۰ ثانیه
    String cached = cmd.get("product:42");   // خواندن؛ null یعنی miss یا منقضی
}
client.shutdown();
نکته‌ی لایسنس که باید بدانی (به‌روز نگه‌دار)

تاریخچه‌ی لایسنسِ Redis پرفرازونشیب بوده. Redis در ۲۰۲۴ از لایسنسِ آزادِ BSD خارج شد که باعث شد بنیادِ لینوکس یک فورکِ کاملاً آزاد به نامِ Valkey (تحتِ BSD 3-Clause) بسازد. سپس Redis در نسخه‌ی ۸ عقب‌نشینی کرد و یک مدلِ سه‌لایسنسی معرفی کرد: RSALv2، SSPLv1 و AGPLv3 (که موردِ تأییدِ OSI است). در عمل امروز دو گزینه‌ی رایج داری: Redis (نسخه‌ی ۸ به بعد) و Valkey (فورکِ کاملاً BSD). برای برنامه‌نویسِ جاوا رابطِ برنامه‌نویسی تقریباً یکسان است و کلاینت‌های بالا با هر دو کار می‌کنند؛ اما اگر حساسیتِ حقوقیِ لایسنس داری، این تمایز را در معماری لحاظ کن.


انتزاعِ کش اسپرینگ: کش اعلانی با آنوتیشن

نوشتنِ دستیِ آن الگوی سه‌مرحله‌ای در همه‌جا خسته‌کننده و خطاخیز است. اسپرینگ یک انتزاعِ کش (cache abstraction) می‌دهد که با چند آنوتیشن، کش را به‌شکلِ اعلانی و شفاف روی متدها اعمال می‌کند — بدون اینکه بدنه‌ی متدت آلوده شود.

انتزاعِ کش مثل یک دستیارِ حافظه‌دار جلوی درِ اتاقت

تصور کن یک دستیار جلوی درِ دفترت نشسته. هر بار کسی سؤالی می‌پرسد، دستیار اول در دفترچه‌اش نگاه می‌کند: اگر جواب را قبلاً یادداشت کرده، همان را می‌دهد و اصلاً درِ اتاقِ تو را نمی‌زند (متد اجرا نمی‌شود). اگر نبود، سؤال را به تو می‌دهد، جوابت را می‌گیرد، در دفترچه یادداشت می‌کند و به پرسنده می‌دهد. @Cacheable دقیقاً همین دستیار است — و زیبایی‌اش این است که خودِ تو (متد) اصلاً نمی‌دانی کشی در کار است.

اول انتزاع را روشن کن و یک provider بده. با اسپرینگ‌بوت، وابستگیِ spring-boot-starter-cache را اضافه کن و روی یک کلاسِ کانفیگ @EnableCaching بگذار. اگر Caffeine در classpath باشد، اسپرینگ خودش یک CaffeineCacheManager می‌سازد؛ اگر Redis باشد، RedisCacheManager.

@Configuration
@EnableCaching
public class CacheConfig {

    @Bean
    public CacheManager cacheManager() {
        CaffeineCacheManager manager = new CaffeineCacheManager("products", "users");
        manager.setCaffeine(Caffeine.newBuilder()
            .maximumSize(10_000)
            .expireAfterWrite(Duration.ofMinutes(5))
            .recordStats());
        return manager;
    }
}

@Cacheable — خواندنِ کش‌شده

@Service
public class ProductService {

    @Cacheable(cacheNames = "products", key = "#id")
    public Product getProduct(long id) {
        // این بدنه فقط در miss اجرا می‌شود
        return productRepository.findById(id);
    }
}

بارِ اول با id = 42 بدنه اجرا و نتیجه کش می‌شود؛ فراخوانی‌های بعدی با همان کلید، بدنه را اصلاً اجرا نمی‌کنند و مستقیم از کش برمی‌گردند.

دامِ کلاسیک: پروکسیِ خودی (self-invocation)

انتزاعِ کش اسپرینگ با پروکسی کار می‌کند: اسپرینگ دورِ bean تو یک لایه می‌کشد که کش را چک کند. اگر یک متد از همان کلاس متدِ @Cacheable دیگری را با this.getProduct(...) صدا بزند، فراخوانی از پروکسی رد نمی‌شود و کش کاملاً دور زده می‌شود — بدنه هر بار اجرا می‌شود انگار کشی نیست. این یکی از پرتکرارترین باگ‌های خاموش است. راه‌حل: فراخوانی را از یک bean دیگر انجام بده، یا آن متد را به یک سرویسِ جدا منتقل کن.

@CacheEvict و @CachePut — بی‌اعتبارسازی و به‌روزرسانی

  • @CacheEvict ورودی را از کش پاک می‌کند — وقتی داده عوض یا حذف شد.
  • @CachePut بدنه را همیشه اجرا می‌کند و نتیجه‌اش را در کش می‌گذارد (write-through اعلانی).
@CachePut(cacheNames = "products", key = "#product.id")
public Product update(Product product) {
    return productRepository.save(product);   // هم DB به‌روز می‌شود هم کش
}

@CacheEvict(cacheNames = "products", key = "#id")
public void delete(long id) {
    productRepository.deleteById(id);
}

@CacheEvict(cacheNames = "products", allEntries = true)
public void reloadAll() {
    // کلِ کشِ products را پاک می‌کند
}

آنوتیشن‌های دیگر: @Caching برای ترکیبِ چند عملیات روی یک متد، و @CacheConfig برای گذاشتنِ تنظیماتِ مشترک (مثل نامِ کش) در سطحِ کلاس. همچنین condition (کش کن فقط اگر شرط برقرار باشد) و unless (کش نکن اگر شرط برقرار باشد، که روی نتیجه ارزیابی می‌شود) کنترلِ ظریف می‌دهند:

@Cacheable(cacheNames = "products", key = "#id",
           unless = "#result == null")   // نتیجه‌ی null را کش نکن
public Product getProduct(long id) { ... }
پرچمِ sync — سپرِ داخلیِ اسپرینگ در برابرِ طوفان

پیش‌فرضِ @Cacheable در برابرِ درخواست‌های هم‌زمانِ miss محافظت نمی‌کند: اگر هزار thread هم‌زمان کلیدِ سردی را بخواهند، هر هزار تا بدنه را اجرا می‌کنند. با @Cacheable(sync = true) اسپرینگ تضمین می‌کند برای یک کلید فقط یک thread بدنه را اجرا کند و بقیه منتظرِ نتیجه‌اش بمانند — همان ادغامِ درخواست. این ساده‌ترین دفاعِ درون‌JVMی در برابرِ طوفانِ کش است. یادت باشد این فقط درونِ یک نمونه کار می‌کند، نه بین نمونه‌ها.


معماریِ دو‌سطحی: L1 (Caffeine) + L2 (Redis)

در سیستم‌های جدی، اغلب هر دو دنیا را با هم می‌خواهی: سرعتِ نانوثانیه‌ایِ محلی و اشتراکِ بین‌نمونه‌ای. این می‌شود کشِ چندسطحی یا near cache.

L1 مثل جیب، L2 مثل کیف، دیتابیس مثل خانه

پول را سه‌جا نگه می‌داری. جیبت (L1 / Caffeine) سریع‌ترین دسترسی را دارد ولی کم جا. کیفت (L2 / Redis) کندتر ولی جادارتر و مشترک با بقیه. و خانه (دیتابیس) منبعِ نهایی. وقتی چیزی می‌خواهی: اول جیب، بعد کیف، بعد خانه — و هر بار که از عمق آوردی، در لایه‌های بالاتر هم می‌گذاری تا دفعه‌ی بعد سریع‌تر باشد. این دقیقاً جریانِ خواندن در کشِ دو‌سطحی است.

public Product get(long id) {
    // L1: کش محلیِ فوق‌سریع
    Product p = caffeine.getIfPresent(id);
    if (p != null) return p;

    // L2: کش مشترکِ Redis
    p = readFromRedis(id);
    if (p != null) {
        caffeine.put(id, p);   // به L1 هم برگردان
        return p;
    }

    // منبعِ اصلی
    p = productRepository.findById(id);
    if (p != null) {
        writeToRedis(id, p);   // پر کردنِ L2
        caffeine.put(id, p);   // و L1
    }
    return p;
}

اما دو‌سطحی یک مشکلِ جدی می‌آورد: بی‌اعتبارسازیِ L1. وقتی داده عوض می‌شود، Redis (مشترک) را می‌شود مستقیم پاک کرد، اما هر نمونه یک Caffeineِ محلیِ مستقل دارد که خبر ندارد. راه‌حلِ رایج: از pub/sub رِدیس استفاده کن — وقتی داده‌ای عوض شد، یک پیام روی یک کانالِ Redis منتشر کن؛ همه‌ی نمونه‌ها مشترکِ آن کانال‌اند و با دریافتِ پیام، ورودیِ متناظر را از Caffeineِ محلیِ خودشان پاک می‌کنند.

near cache در کتابخانه‌های آماده

لازم نیست همیشه این را دستی بسازی. Redisson یک RLocalCachedMap می‌دهد که دقیقاً همین near cache را با بی‌اعتبارسازیِ خودکارِ مبتنی‌بر pub/sub پیاده کرده. اما فهمیدنِ مکانیزم (L1 محلی + L2 مشترک + کانالِ بی‌اعتبارسازی) مهم‌تر از دانستنِ نامِ کلاس است — چون در مصاحبه همین مکانیزم را می‌پرسند.


طوفانِ کش: thundering herd و مهارش

حالا به یکی از خطرناک‌ترین پدیده‌های کش می‌رسیم که بارها اشاره‌اش کردیم. اسم‌هایش زیاد است: cache stampede، thundering herd، dogpile.

طوفانِ کش مثل هجومِ همزمانِ همه به یک در

تصور کن یک کلیدِ خیلی داغ — مثلاً قیمتِ صفحه‌ی اولِ فروشگاه که ثانیه‌ای هزار بار خوانده می‌شود — ناگهان TTL‌اش تمام می‌شود و از کش می‌رود. در همان لحظه، هزار درخواست که دنبالِ آن کلید بودند، هم‌زمان miss می‌بینند و هم‌زمان به دیتابیس هجوم می‌برند تا بازش سازند. دیتابیسی که راحت هزار درخواستِ کش‌شده را تاب می‌آورد، حالا زیرِ هزار کوئریِ سنگینِ هم‌زمان زانو می‌زند — و گاهی کلِ سیستم را با خودش پایین می‌کشد. بدترین قسمت؟ همه‌ی آن هزار تا دارند دقیقاً همان کار را انجام می‌دهند.

چند راهِ اثبات‌شده برای مهارش:

۱) ادغامِ درخواست / single-flight. همان چیزی که LoadingCache و @Cacheable(sync=true) می‌دهند: تضمین کن برای یک کلید فقط یک محاسبه اجرا شود و بقیه پشتش صف بکشند. ساده‌ترین و اولین دفاع — اما فقط درونِ یک JVM.

۲) قفلِ توزیع‌شده (mutex). برای طوفانِ بین‌نمونه‌ای، فقط اجازه بده یک نمونه در کلِ کلاستر کلید را بازسازی کند؛ بقیه یا کمی صبر می‌کنند یا داده‌ی بیاتِ قبلی را می‌گیرند. با Redis این را با SET key value NX PX <ms> می‌سازی — یک قفلِ اتمیکِ زمان‌دار:

// فقط یک نمونه قفل را می‌گیرد و بازسازی می‌کند
boolean gotLock = "OK".equals(
    cmd.set("lock:product:42", token, SetArgs.Builder.nx().px(3000)));
if (gotLock) {
    try {
        Product fresh = productRepository.findById(42L);
        writeToRedis(42L, fresh);
    } finally {
        // آزادسازیِ امنِ قفل (فقط اگر تو صاحبش باشی) با اسکریپت Lua
        releaseLockIfOwner("lock:product:42", token);
    }
}

۳) انقضای احتمالاتیِ زودهنگام (XFetch). یک تکنیکِ ظریف و زیبا: به‌جای اینکه همه سرِ یک لحظه‌ی دقیق منقضی شوند، هرچه به انقضا نزدیک‌تر می‌شوی، هر درخواست با احتمالِ فزاینده‌ای تصمیم می‌گیرد کش را زودتر (در پس‌زمینه) بازسازی کند. چون این تصمیم مستقل و تصادفی است، فقط یکی از هزار درخواست معمولاً زودتر بازسازی می‌کند و بقیه هنوز کشِ معتبر را می‌گیرند. فرمولِ آکادمیک: بازسازی کن اگر now - delta * beta * ln(random()) >= expiry، که delta زمانِ محاسبه و beta (پیش‌فرض ۱٫۰) پارامترِ تنظیم است.

۴) پراکندگیِ TTL (jitter). اگر هزار کلید را با هم و با TTLِ دقیقاً یکسان پر کنی، همه‌شان هم‌زمان منقضی می‌شوند و طوفانِ دسته‌جمعی می‌سازند. راه‌حلِ ساده: به هر TTL کمی تصادفی‌بودن اضافه کن (مثلاً ۵ دقیقه ± ۳۰ ثانیه) تا انقضاها پخش شوند.

۵) stale-while-revalidate. داده‌ی بیات را نگه دار و فوراً همان را بده، و در پس‌زمینه تازه‌سازی کن. این دقیقاً همان کاری است که refreshAfterWrite در Caffeine می‌کند.

کدام راه را انتخاب کنی؟

برای طوفانِ درونِ یک JVM، LoadingCache یا sync=true معمولاً کافی است — ساده و رایگان. برای طوفانِ بین‌نمونه‌ای روی یک کلیدِ داغ و گران، قفلِ توزیع‌شده‌ی Redis یا XFetch لازم است. و jitter را تقریباً همیشه اعمال کن؛ ارزان است و از طوفان‌های دسته‌جمعیِ ناشی از انقضای هم‌زمان جلوگیری می‌کند. این‌ها هم‌دیگر را رد نمی‌کنند؛ اغلب چندتا را با هم به کار می‌بری.

سه دیوِ کش را با هم اشتباه نگیر

سه مشکلِ متفاوت با اسم‌های شبیه وجود دارد. Cache stampede (که گفتیم): هجومِ هم‌زمان بعد از منقضی‌شدنِ یک کلیدِ داغ. Cache penetration (نفوذ): درخواستِ مکرر برای کلیدی که اصلاً وجود ندارد (مثلاً id = -1)، پس هر بار miss می‌شود و مستقیم به دیتابیس می‌رود — راهِ حل: نتیجه‌ی null/خالی را هم با TTLِ کوتاه کش کن، یا Bloom filter بگذار. Cache avalanche (بهمن): وقتی بخشِ بزرگی از کش هم‌زمان منقضی یا خودِ Redis کرش می‌کند و ناگهان همه‌ی بار به دیتابیس می‌ریزد — راهِ حل: jitter، افزونگیِ Redis، و circuit breaker. در مصاحبه تفکیکِ این سه امتیازِ بزرگی است.


بی‌اعتبارسازی و مبادله‌ی سازگاری

رسیدیم به همان «سخت‌ترین مسئله». وقتی داده در منبع عوض می‌شود، کش چطور بفهمد که نسخه‌اش کهنه شده؟ چند استراتژی:

  • مبتنی‌بر TTL (منقضی‌شونده): ساده‌ترین. هر ورودی بعد از مدتی خودبه‌خود می‌میرد. سازگاری «سرانجام‌گرا» است: پنجره‌ی بیاتی حداکثر به‌اندازه‌ی TTL است. برای اکثرِ داده‌ها همین کافی است و شگفت‌انگیز مؤثر است.
  • مبتنی‌بر رویداد (صریح): وقتی داده عوض شد، فعالانه ورودیِ کش را پاک یا به‌روز کن (@CacheEvict/@CachePut). دقیق‌تر است اما شکننده‌تر: باید همه‌ی مسیرهای تغییرِ داده را پیدا کنی و هیچ‌کدام را جا نیندازی — کاری که در عمل سخت است.
  • کلیدِ نسخه‌دار (versioned key): به‌جای پاک‌کردن، کلید را عوض کن. مثلاً product:42:v7؛ وقتی محصول عوض شد، نسخه را به v8 ببر. حالا خواننده‌ها که دنبالِ v8 می‌گردند خودبه‌خود miss می‌بینند و کلیدِ قدیم بی‌استفاده منقضی می‌شود. ظریف و بی‌نیاز از هماهنگیِ حذف.
چرا invalidation این‌قدر سخت است

مشکلِ عمیقِ بی‌اعتبارسازیِ صریح این است که یک مسئله‌ی سازگاریِ توزیع‌شده است. بین «نوشتن در دیتابیس» و «پاک‌کردنِ کش» یک شکاف زمانی هست، و در آن شکاف — یا اگر پاک‌کردن شکست بخورد — کش داده‌ی غلط می‌دهد. حتی ترتیب مهم است: اگر اول کش را پاک کنی و بعد دیتابیس را بنویسی، یک خواننده‌ی هم‌زمان می‌تواند مقدارِ قدیمِ دیتابیس را دوباره در کش بگذارد و آن را برای همیشه بیات کند. برای همین بسیاری از سیستم‌های باتجربه به‌جای تلاش برای سازگاریِ کامل، روی TTLِ کوتاه تکیه می‌کنند: به‌جای «همیشه درست»، «حداکثر N ثانیه غلط» را می‌پذیرند، چون کنترلش بسیار ساده‌تر است.

مبادله‌ی سازگاری مثل روزنامه در برابر تابلوی زنده‌ی فرودگاه

دو مدلِ ذهنی برای تازگی داری. روزنامه (کش با TTL): یک عکسِ فوریِ لحظه‌ی چاپ است؛ می‌دانی ممکن است تا شبِ همان روز کمی کهنه شود، اما ارزان و سریع در دسترس است و برای اکثرِ کارها کافی. تابلوی زنده‌ی پروازها (بدونِ کش، یا write-through): همیشه لحظه‌ای و دقیق است، اما هزینه‌ی نگه‌داشتنش بسیار بالاست. سؤالِ مهندسیِ درست این نیست «کدام بهتر است»، بلکه «چقدر کهنگی برای این داده قابل‌قبول است؟» قیمتِ سهام؟ ثانیه‌ای. نامِ دسته‌بندیِ محصول؟ ساعتی هم اشکالی ندارد.

نکته‌ی معماریِ نهایی: کش تقریباً همیشه یعنی پذیرفتنِ سازگاریِ سرانجام‌گرا (eventual consistency). لحظه‌ای که یک کپی نگه می‌داری، پنجره‌ای از ناهم‌خوانی می‌سازی. کارِ مهندسِ خوب حذفِ این پنجره نیست (که اغلب ناممکن یا بسیار گران است)، بلکه کوچک‌کردنِ آگاهانه‌ی آن تا حدی که برای دامنه‌ی مسئله بی‌ضرر باشد.


دام‌ها و بهترین‌شیوه‌ها

بیا مهم‌ترین تله‌ها را یک‌جا جمع کنیم — همان‌هایی که در کدِ واقعی یا مصاحبه گازت می‌گیرند:

  • کشِ بی‌سقف = نشتِ حافظه. همیشه maximumSize/maximumWeight (یا TTL) بگذار؛ کشِ نامحدود دیر یا زود OutOfMemoryError می‌دهد.
  • پروکسیِ خودی، کش را دور می‌زند. فراخوانیِ this.method() درونِ همان bean، لایه‌ی کشِ اسپرینگ را رد نمی‌کند.
  • کش‌کردنِ null بدونِ فکر. اگر null را بی‌کنترل کش کنی و منبع بعداً مقدار بگیرد، بیاتی می‌ماند؛ اگر اصلاً کش نکنی، در برابرِ نفوذ (penetration) آسیب‌پذیری. تصمیمِ آگاهانه بگیر (اغلب: null را با TTLِ کوتاه کش کن).
  • TTLِ یکسان برای همه = بهمن. jitter اضافه کن تا انقضاها پخش شوند.
  • کش‌کردنِ آبجکتِ تغییرپذیر و اشتراکش. اگر آبجکتی که کش کرده‌ای تغییرپذیر باشد و فراخواننده تغییرش دهد، نسخه‌ی کش‌شده هم عوض می‌شود (در کش محلی). آبجکت‌های تغییرناپذیر کش کن یا کپی بده.
  • کلیدِ سریال‌سازی‌شده‌ی ناپایدار در Redis. اگر آبجکت جاوا را با سریال‌سازیِ بومیِ جاوا در Redis بگذاری، تغییرِ کلاس می‌تواند دیسریالایز را بشکند و — بدتر — سطحِ حمله باز کند. از JSON یا یک قالبِ نسخه‌دارِ صریح استفاده کن.
  • فراموشیِ پایش. بدونِ hit ratio نمی‌دانی کشت کار می‌کند. recordStats و متریک‌ها را جدی بگیر.

بهترین‌شیوه‌ها:

  • برای هر کش، این سه سؤال را جواب بده: سقفِ اندازه چقدر؟ سیاستِ انقضا چیست (TTL یا TTI و چند)؟ نرخِ اصابتم را چطور می‌بینم؟
  • کشِ محلی برای داده‌ی فقط-خواندنیِ پرتکرار؛ Redis برای اشتراکِ بین‌نمونه‌ای؛ دو‌سطحی برای هر دو.
  • برای طوفان، از ادغامِ درخواست شروع کن (sync/LoadingCache)، و برای کلیدهای داغِ بین‌نمونه‌ای قفلِ توزیع‌شده یا XFetch اضافه کن.
  • به‌جای شکارِ سازگاریِ کامل، TTLِ کوتاه بپذیر مگر جایی که واقعاً بیاتی غیرقابل‌تحمل است.

پرسش‌های مصاحبه

حالا همه‌چیز را در قالبِ سؤال‌های واقعیِ مصاحبه‌ی سنیور تمرین کنیم. اول خودت جواب بده، بعد پاسخ را باز کن.

۱) تفاوتِ cache-aside و read-through دقیقاً چیست؟

در cache-aside خودِ کدِ برنامه مسئولِ همه‌چیز است: کش را می‌خواند، اگر miss بود منبع را می‌خواند، و خودش نتیجه را در کش می‌گذارد. کش صرفاً یک انبارِ منفعلِ کلید-مقدار است. در read-through این منطق به کش سپرده می‌شود: تو یک تابعِ لود می‌دهی و کش خودش هنگام miss آن را صدا می‌زند و پر می‌شود. تفاوتِ عملیِ مهم: read-through معمولاً درخواست‌های هم‌زمان روی یک کلید را ادغام (coalesce) می‌کند، پس در برابرِ طوفان امن‌تر است؛ cache-aside هیچ محافظتِ خودکاری ندارد. LoadingCache در Caffeine نمونه‌ی read-through است.

۲) TTL و TTI چه فرقی دارند و در Caffeine هرکدام کدام متد است؟

TTL (Time To Live) یعنی ورودی حداکثر مدتِ مشخصی از لحظه‌ی نوشتن زنده می‌ماند، فارغ از تعدادِ دسترسی‌ها — در Caffeine یعنی expireAfterWrite. مناسبِ داده‌ای که در منبع کهنه می‌شود (نرخ ارز، قیمت). TTI (Time To Idle) یعنی ورودی تا وقتی استفاده می‌شود زنده می‌ماند و فقط بی‌مصرف‌ماندنِ طولانی می‌کُشدش؛ هر دسترسی ساعت را ریست می‌کند — در Caffeine یعنی expireAfterAccess. مناسبِ نشستِ کاربر یا داده‌ی موقت. می‌شود هر دو را با هم گذاشت.

۳) (ظریف) تفاوتِ expireAfterWrite و refreshAfterWrite چیست؟

expireAfterWrite سخت است: بعد از موعد ورودی حذف می‌شود و درخواستِ بعدی منتظرِ لودِ تازه می‌ماند (یعنی یک miss با تأخیر روی مسیرِ اصلی). refreshAfterWrite نرم است: بعد از موعد، مقدارِ کهنه هنوز هست و فوراً برگردانده می‌شود، اما یک لودِ تازه در پس‌زمینه راه می‌افتد تا دفعه‌ی بعد به‌روز باشد. یعنی refresh تأخیرِ کاربر را حذف می‌کند به قیمتِ یک بیاتیِ لحظه‌ای؛ expire تازگی را تضمین می‌کند به قیمتِ تأخیر. اغلب هر دو را با هم می‌گذاری، با refresh کوتاه‌تر از expire.

۴) W-TinyLFU چیست و چرا از LRU و LFU بهتر است؟

W-TinyLFU الگوریتمِ حذفِ Caffeine است که بسامد و اخیربودن را ترکیب می‌کند. مشکلِ LRU این است که در برابرِ پویشِ ترتیبی فرومی‌پاشد (داده‌ی یک‌بارمصرف داغ‌ها را بیرون می‌ریزد). مشکلِ LFU حافظه‌ی زیاد برای شمارنده‌ها و گیرکردن به داغیِ گذشته است. W-TinyLFU هر دو را حل می‌کند: یک frequency sketch (شبیه Count-Min Sketch) بسامد را با چند بیت تخمین می‌زند؛ یک پنجره‌ی ورودی به تازه‌واردها فرصت می‌دهد؛ یک دربانِ پذیرش فقط ورودی‌ای را می‌پذیرد که بسامدِ تخمینی‌اش از قربانی بیشتر باشد (پویش را خنثی می‌کند)؛ و پیرسازیِ شمارنده‌ها داغیِ کهنه را فراموش می‌کند. نتیجه: نرخِ اصابتی نزدیک به بهینه با سربارِ حافظه‌ی ناچیز.

۵) کِی کش محلی (Caffeine) و کِی توزیع‌شده (Redis)؟

Caffeine وقتی که داده فقط-خواندنیِ پرتکرار است، تحملِ ناهم‌خوانیِ کوتاه بین نمونه‌ها را داری، و بیشترین سرعت (نانوثانیه، بدونِ شبکه و سریال‌سازی) را می‌خواهی. Redis وقتی که چند نمونه باید یک نسخه‌ی مشترک و هماهنگ ببینند، یا کش باید از کرشِ برنامه جان به‌در ببرد، یا حجمِ داده از heap بزرگ‌تر است. در سیستم‌های جدی اغلب هر دو با هم: Caffeine به‌عنوان L1 و Redis به‌عنوان L2 (near cache). جمله‌ی طلایی: «محلی برای سرعت، توزیع‌شده برای هماهنگی.»

۶) طوفانِ کش (thundering herd) چیست و چطور مهارش می‌کنی؟

وقتی یک کلیدِ داغ منقضی می‌شود و هزاران درخواستِ هم‌زمان با هم miss می‌بینند و با هم به دیتابیس هجوم می‌برند تا بازش سازند، دیتابیس زیرِ بارِ کارِ تکراری زانو می‌زند. راه‌های مهار: (۱) ادغامِ درخواست/single-flight با LoadingCache یا @Cacheable(sync=true) — یک محاسبه به‌ازای کلید، درونِ یک JVM. (۲) قفلِ توزیع‌شده با Redis SET NX PX برای طوفانِ بین‌نمونه‌ای. (۳) انقضای احتمالاتیِ زودهنگام (XFetch) که به‌صورتِ تصادفی فقط یکی زودتر بازسازی کند. (۴) jitter روی TTL تا انقضاها پخش شوند. (۵) stale-while-revalidate (مثلِ refreshAfterWrite). معمولاً چندتا را با هم به کار می‌بری.

۷) (تمایز) stampede، penetration و avalanche را از هم جدا کن.

Stampede/thundering herd: هجومِ هم‌زمان بعد از منقضی‌شدنِ یک کلیدِ داغ. راه‌حل: ادغامِ درخواست، قفل، XFetch. Penetration (نفوذ): درخواستِ مکرر برای کلیدی که اصلاً وجود ندارد، پس هر بار miss و مستقیم به دیتابیس. راه‌حل: کش‌کردنِ نتیجه‌ی null با TTLِ کوتاه، یا Bloom filter. Avalanche (بهمن): منقضی‌شدنِ بخشِ بزرگی از کش هم‌زمان (یا کرشِ خودِ Redis) که یک‌باره همه‌ی بار را روی دیتابیس می‌ریزد. راه‌حل: jitter، افزونگیِ Redis، circuit breaker. تفکیکِ دقیقِ این سه، نشانه‌ی سنیوریتی است.

۸) چرا @Cacheable گاهی «هیچ‌کاری نمی‌کند» و کش دور زده می‌شود؟

تقریباً همیشه به‌خاطرِ self-invocation. انتزاعِ کش اسپرینگ با پروکسی کار می‌کند؛ لایه‌ی کش فقط وقتی فعال می‌شود که فراخوانی از بیرونِ bean و از طریقِ پروکسی بیاید. اگر متدی از همان کلاس، متدِ @Cacheable دیگری را با this.method(...) صدا بزند، از پروکسی رد نمی‌شود و کش کاملاً دور زده می‌شود — بدنه هر بار اجرا می‌شود. راه‌حل: فراخوانی را از یک bean دیگر انجام بده یا متد را به سرویسِ جدا منتقل کن. دلایلِ دیگرِ محتمل: نبودِ @EnableCaching، یا شرطِ condition/unless که مانع کش شده.

۹) استراتژی‌های بی‌اعتبارسازی کدام‌اند و کدام را ترجیح می‌دهی؟

سه استراتژیِ اصلی: TTL (هر ورودی خودبه‌خود می‌میرد؛ ساده و مقاوم، اما بیاتی تا سقفِ TTL)، مبتنی‌بر رویداد (با @CacheEvict/@CachePut صریحاً پاک/به‌روز کن؛ دقیق اما شکننده چون باید همه‌ی مسیرهای تغییر را بپوشانی)، و کلیدِ نسخه‌دار (به‌جای پاک‌کردن، نسخه‌ی کلید را بالا ببر تا خواننده‌ها خودبه‌خود miss ببینند). در عمل اغلب TTLِ کوتاه را ترجیح می‌دهم چون بی‌اعتبارسازیِ صریح یک مسئله‌ی سازگاریِ توزیع‌شده است (شکاف بین نوشتنِ DB و پاک‌کردنِ کش، و خطرِ شکستِ پاک‌کردن)؛ می‌شود TTL و رویداد را هم ترکیب کرد.

۱۰) (ظریف ترتیب) اگر بخواهی دیتابیس را بنویسی و کش را پاک کنی، اول کدام؟

اول دیتابیس را بنویس، بعد کش را پاک کن (الگوی cache-aside invalidation). اگر برعکس عمل کنی — اول کش را پاک کنی بعد DB را بنویسی — یک خواننده‌ی هم‌زمان می‌تواند بینِ دو عمل، مقدارِ قدیمِ دیتابیس را بخواند و دوباره در کش بگذارد، و آن را برای همیشه بیات کند. حتی «نوشتن در DB بعد پاک‌کردنِ کش» هم بی‌نقص نیست (اگر پاک‌کردن شکست بخورد کش بیات می‌مانَد)، برای همین معمولاً یک TTL هم به‌عنوانِ تورِ ایمنی می‌گذاریم. این چراییِ محبوبیتِ «invalidate به‌جای update» هم هست: پاک‌کردن ساده‌تر از هم‌گام‌نگه‌داشتن است.

۱۱) چرا Redis را «تک‌رشته‌ای» می‌گویند و چه پیامدی دارد؟

منظور این است که پردازشِ فرمان‌ها روی یک رشته‌ی منطقی سریال می‌شود (نسخه‌های جدید I/O شبکه را چندرشته‌ای کرده‌اند، اما اجرای خودِ فرمان همچنان سریال است). پیامدِ مثبت: هر فرمان به‌صورتِ طبیعی اتمیک است بدونِ قفلِ پیچیده، پس دستوری مثل INCR یا SET NX بی‌هیچ race condition کار می‌کند — که همان چیزی است که قفلِ توزیع‌شده را ممکن می‌کند. پیامدِ منفی: یک فرمانِ سنگین (مثلِ KEYS * روی دیتابیسِ بزرگ) کلِ سرور را بلاک می‌کند، پس باید از فرمان‌های کند پرهیز کرد. سرعتِ کلی از درون‌حافظه‌ای‌بودن و همین سادگی می‌آید.

۱۲) (تله‌ی حافظه) چرا کش‌کردنِ یک آبجکتِ تغییرپذیر در Caffeine خطرناک است؟

چون کشِ محلی ارجاعِ آبجکت را نگه می‌دارد، نه یک کپی. اگر همان آبجکتِ کش‌شده را به فراخواننده بدهی و او فیلدی از آن را عوض کند، نسخه‌ی داخلِ کش هم عوض می‌شود — چون هر دو به یک آبجکت اشاره می‌کنند. نتیجه: کش داده‌ی دستکاری‌شده و ناخواسته می‌دهد و باگ‌های به‌شدت گیج‌کننده می‌سازد. راه‌حل: آبجکت‌های تغییرناپذیر کش کن (مثلِ record با کپیِ دفاعی)، یا هنگامِ خواندن/نوشتن کپی بگیر. در Redis این مشکل نیست چون داده سریال‌سازی و کپی می‌شود، اما آن‌جا هزینه‌ی سریال‌سازی را می‌دهی.

۱۳) (معماری) کش دو‌سطحی چطور کار می‌کند و بزرگ‌ترین چالشش چیست؟

L1 یک کش محلیِ سریع (Caffeine) درونِ هر نمونه، و L2 یک کش مشترک (Redis) است. خواندن: اول L1، بعد L2، بعد دیتابیس؛ و هر بار که از عمق آوردی، لایه‌های بالاتر را هم پر می‌کنی. مزیت: سرعتِ محلی + اشتراکِ بین‌نمونه‌ای. بزرگ‌ترین چالش بی‌اعتبارسازیِ L1 است: وقتی داده عوض می‌شود، L2 (مشترک) را می‌شود مستقیم پاک کرد اما هر نمونه یک L1ِ مستقل دارد که خبر ندارد. راه‌حلِ رایج: pub/sub رِدیس — با تغییرِ داده یک پیام منتشر کن و همه‌ی نمونه‌ها ورودیِ متناظر را از L1ِ خودشان پاک کنند. کتابخانه‌هایی مثل Redisson (RLocalCachedMap) این را آماده دارند.

۱۴) کش‌زدن چه هزینه‌های پنهانی دارد و کِی *نباید* کش زد؟

هزینه‌های پنهان: (۱) ناسازگاری — لحظه‌ای که کپی نگه می‌داری پنجره‌ی بیاتی می‌سازی. (۲) پیچیدگیِ عملیاتی — یک سیستمِ حالت‌دارِ دیگر که باید پایش، اندازه‌گیری و دیباگ شود. (۳) مصرفِ حافظه و خطرِ نشت اگر بی‌سقف باشد. (۴) هزینه‌ی سریال‌سازی در کشِ توزیع‌شده. نباید کش زد وقتی: داده به‌سرعت عوض می‌شود و بیاتی غیرقابل‌قبول است؛ نرخِ اصابت پایین است (کشِ سرد فقط حافظه و پیچیدگی هدر می‌دهد)؛ منبعِ اصلی خودش به‌قدرِ کافی سریع و ارزان است؛ یا داده حساس است و نگه‌داشتنِ کپی‌اش ریسکِ امنیتی دارد. کش یک ابزار است، نه یک واجب — اول hit ratio را اندازه بگیر، بعد قضاوت کن.


جمع‌بندی
  • چرا کش: سریع‌ترین کار، کارِ نکرده است. کش هم تأخیر را کم می‌کند و هم بارِ منبعِ اصلی را — گاهی هدفِ اصلی، محافظت از دیتابیس در برابرِ فروپاشی است.
  • الگوها: cache-aside (دستی، ساده)، read-through (کش با تابعِ لود و ادغامِ خودکار)، write-through (تازه ولی نوشتنِ کند)، write-behind (سریع ولی خطرِ ازدست‌رفتنِ داده)، refresh-ahead (تازه‌سازیِ پیش‌دستانه).
  • حذف: LRU در برابرِ پویش می‌شکند، LFU به گذشته گیر می‌کند؛ W-TinyLFU در Caffeine با frequency sketch + پنجره + دربانِ پذیرش + پیرسازی، بهترینِ هر دو را می‌گیرد.
  • Caffeine: کشِ محلیِ نانوثانیه‌ای؛ expireAfterWrite (TTL) در برابر expireAfterAccess (TTI)، و refreshAfterWrite که بیاتِ فوری می‌دهد و در پس‌زمینه تازه می‌کند. همیشه maximumSize و recordStats.
  • Redis: کشِ مشترکِ بیرونی برای هماهنگیِ بین‌نمونه‌ای و دوامِ فراتر از کرش؛ کلاینت‌ها: Lettuce (پیش‌فرض)، Jedis، Redisson. اجرای فرمانِ سریال ⇒ اتمیک‌بودنِ طبیعی.
  • اسپرینگ: @Cacheable/@CacheEvict/@CachePut کشِ اعلانی می‌دهند؛ مراقبِ self-invocation باش؛ sync=true سپرِ درون‌JVMی در برابرِ طوفان است.
  • طوفان و دوستانش: stampede (ادغام/قفل/XFetch)، penetration (کشِ null/Bloom filter)، avalanche (jitter/افزونگی). jitter را تقریباً همیشه اعمال کن.
  • سازگاری: کش یعنی پذیرفتنِ سازگاریِ سرانجام‌گرا. اغلب TTLِ کوتاه بهتر از تعقیبِ سازگاریِ کاملِ شکننده است. اول در DB بنویس، بعد کش را پاک کن.

Let's accept one simple truth up front: the fastest thing an application can do is the work it never does. If you already have the answer to a question, you don't need to hit the database again, recompute it again, or ask a remote service again. A cache is exactly that: a small, close-at-hand notebook where you keep frequent answers so that next time you can hand them over in nanoseconds instead of milliseconds.

In this chapter you won't just learn to slap @Cacheable on a method and move on. You'll understand what pain a cache cures, why it sometimes becomes the source of the worst bugs, what the two reference tools of the Java world — Caffeine for local caching and Redis for distributed caching — actually do, when to pick which, and how to talk about the hard trade-offs of caching in a senior interview.

Roadmap for this chapter

The path we'll walk together:

  1. Part 0 — base vocabulary: hit, miss, eviction, TTL. Analogy first, technical term second.
  2. Why cache — the latency numbers and why caching makes the difference between "slow" and "fast."
  3. Cache patterns — cache-aside, read-through, write-through, write-behind, refresh-ahead (with a comparison table).
  4. Eviction policies — LRU, LFU, and why Caffeine chose the modern W-TinyLFU.
  5. Caffeine in practice — building a cache, LoadingCache, expireAfterWrite vs expireAfterAccess, and refreshAfterWrite.
  6. Local vs distributed — Caffeine vs Redis, and when to use which.
  7. Redis and its clients — Lettuce, Jedis, Redisson, and an important licensing note.
  8. Spring's cache abstraction@Cacheable, @CacheEvict, @CachePut, and the life-saving sync flag.
  9. Two-tier architecture, cache stampede and taming it, invalidation, and the consistency trade-off.
  10. Pitfalls, best practices, and interview questions with full answers.

Part 0 — a few words you must feel before any code

Before we get to code, a few terms recur throughout this chapter. Let me plant them with analogies right now.

A cache is like a chef's prep counter

A professional chef doesn't fetch every ingredient from the cold storage (the database); they keep the most-used ones — salt, oil, chopped onions — on the prep counter right beside them (the cache). A trip to cold storage takes seconds; reaching to the counter is instant. But the counter can't hold the entire pantry, so the chef constantly decides what stays on the counter and what goes back. That's the whole story of caching: a small, expensive, fast space, and the craft of deciding what deserves to stay on it.

A few words you'll see relentlessly from here on:

  • Hit: the answer you wanted was in the cache and handed straight back. Like finding the salt on the counter.
  • Miss: the answer wasn't in the cache, so you had to go to the real source (database, service, computation). Like a trip to cold storage.
  • Hit ratio: the percentage of requests that were hits. A cache with a 95% hit ratio means 95 of every 100 requests never even reached the database. This is the single most important health number of a cache.
  • Eviction: when the cache fills up, it must throw something out to make room for a newcomer. Which one it throws out is the "eviction policy" — the soul of this chapter.
  • TTL — Time To Live: how long, at most, an entry is allowed to stay in the cache, regardless of how often it's read. Like the expiry date on a carton of milk.
  • Evict vs expire: "expire" means its time ran out (TTL); "evict" means room ran out and we pushed it out. Two different reasons to leave.
  • Stale: data that's still in the cache but no longer matches the source — like a price that's 100 in the cache but has become 120 in the database. The core trade-off of caching is always with this "staleness."
The two hard problems of computer science

A famous joke in software engineering goes: "There are only two hard problems in computer science: cache invalidation and naming things." It's a joke, but the bottom of it is serious: adding a cache is easy; making sure the cache never serves wrong data is one of the hardest jobs in engineering. The entire second half of this chapter is about that hardness.


Why cache at all? The story of the numbers

Caching isn't a decorative optimization; it comes from a physical reality: not all memory is equally fast. Between reading from the application's local memory and going to a database over the network, there are several orders of magnitude of difference.

Operation Approx. latency Human-scale analogy
Read from local in-process cache (Caffeine) ~100 ns Picking something off your desk
Round trip to Redis on the same network ~0.5–1 ms Asking the colleague next door
Simple query to a SQL database ~5–30 ms Going to the archive downstairs
Call to an external service over the internet ~50–500 ms Mailing a letter to another city

The gap between a local cache and a database is roughly fifty-thousand-fold. Now imagine a product page opened a thousand times per second, each time reading the same static product info from the database. Cache it, and a thousand queries per second collapse to nearly zero, and the database breathes.

A cache saves two things at once

First, latency: the user gets their answer faster. Second, load: the primary source (usually the most expensive and fragile part of the system, like the database) is relieved of the pressure of repetitive requests. Very often the real goal of a cache is not user speed but protecting the database from collapse. Remember that distinction in interviews.

But as we said, this speed isn't free. The moment you keep a copy of data somewhere else, you have two versions of the truth — and wherever there are two versions of the truth, there's a chance of divergence. So let's first see how data enters and leaves the cache.


Cache patterns: how data moves

Who is responsible for filling the cache and writing to the database creates different patterns. Confusing these is the root of a lot of muddled thinking.

1) Cache-aside (lazy loading)

The most common pattern, and the one you usually write when caching by hand. Here the application code is responsible for everything; the cache is just a simple key-value store standing beside you (hence "aside").

Cache-aside is like your home fridge

When you're hungry you open the fridge first. If food is there (hit), you eat. If not (miss), you go to the store, buy it, and put a copy in the fridge for next time. The fridge never goes shopping itself; you're the narrator and director of the story. That's exactly cache-aside.

public Product getProduct(long id) {
    // 1) check the cache first
    Product cached = cache.getIfPresent(id);
    if (cached != null) {
        return cached;              // hit
    }
    // 2) miss — go to the real source
    Product fromDb = productRepository.findById(id);
    // 3) put the result back for next time
    if (fromDb != null) {
        cache.put(id, fromDb);
    }
    return fromDb;
}

Its strength is simplicity and total control; its weakness is that you must repeat this "three-step pattern" everywhere by hand, and if you forget, you get a bug. Its bigger problem, which we'll see later: if a thousand requests simultaneously miss, all thousand storm the database at once (that's the cache stampede).

2) Read-through

Here you no longer write those three steps yourself; you tell the cache "if you don't have it, call this function and fill yourself." The cache becomes the read intermediary. Caffeine's LoadingCache is exactly this.

LoadingCache<Long, Product> cache = Caffeine.newBuilder()
    .maximumSize(10_000)
    .build(id -> productRepository.findById(id));  // the "load function"

// now you just say:
Product p = cache.get(42L);   // if absent, the cache calls the loader and stores it

The subtle-but-important difference from cache-aside: the "how to fill" logic lives in one place beside the cache, not scattered across every call site. And more importantly — as we'll see — a loading cache usually coalesces concurrent requests for the same key, meaning the load function runs only once even if a thousand threads want that key at the same time.

3) Write-through

The write-side mirror of read-through: every time you write, you update both the cache and the database synchronously and at the same time. Benefit: the cache never falls behind the database. Cost: every write must wait for the database, so writes get slower.

4) Write-behind / Write-back

Here you write into the cache first and return immediately; the database write happens later, asynchronously and often in batches.

Write-behind is like a store cashier

The cashier takes your money and gives a receipt and off you go (writing to the cache, fast). They collect the cash all day and take it to the bank in one trip at closing (batch write to the database). It's very fast, but there's a risk: if the store burns down midway (a server crash), the cash that hasn't reached the bank is lost. Write-behind is the same: fast, but on a crash it risks losing the queued writes.

5) Refresh-ahead

A pattern whose name doesn't fool you: the cache proactively reloads a "hot" entry in the background before it expires, so the user never hits a miss. Caffeine provides this via refreshAfterWrite, which we'll see shortly.

The two trade-offs that separate these patterns

Keep two axes in mind: data freshness vs write speed, and code simplicity vs automatic protection. Write-through is freshest but slows writes; write-behind has the fastest writes but risks data loss and adds complexity; cache-aside has the simplest mental model but no automatic protection against stampedes; read-through/refresh-ahead give that protection at the cost of handing logic to the cache. The right choice depends on which of these is costlier for you.

Pattern Who writes/reads Freshness Main risk When it fits
Cache-aside App code, by hand Medium Stampede, forgetting to update General use, want full control
Read-through Cache, via load fn Medium Library dependence Read-heavy, want auto-coalescing
Write-through Cache, synchronously High (always fresh) Slow writes Where staleness is unacceptable
Write-behind Cache, deferred Low (temporarily) Data loss on crash Write-heavy, tolerate small loss
Refresh-ahead Cache, proactively in background High for hot keys Wasted loads for cold data Hot, expensive keys

Eviction policies: when the cache fills, which one goes?

A cache is bounded. When it's full and a newcomer arrives, it must sacrifice one. Which one you sacrifice directly determines the hit ratio — and that's what separates Caffeine from a plain HashMap.

LRU — Least Recently Used

LRU is like cleaning out your wardrobe

The simple and popular LRU policy says: "whatever you haven't touched for the longest, goes first." Like your wardrobe: the shirt you haven't worn in a year is the first candidate for the donation bag. The logic is simple and often good: something used recently is likely to be used again soon (the principle of "temporal locality").

But LRU has a famous weakness: it collapses under a sequential scan. Imagine your cache full of hot data, and someone reads the entire table once (say, a heavy report). That scan brings in thousands of one-shot entries, each counted as "recently used," and pushes out all your actual hot data. LRU can't tell that those were requested only once.

LFU — Least Frequently Used

LFU says: "whatever was used least often, goes first" — deciding by count of accesses, not by time of last access. This is more scan-resistant (a one-shot item has frequency 1 and is sacrificed quickly). But LFU has two weaknesses of its own: keeping a counter for everything costs memory, and it gets stuck in the past — something that was very hot last month but is cold now still has a high count and stubbornly stays in the cache.

W-TinyLFU — Caffeine's smart choice

This is where Caffeine shines. Caffeine — originally a rewrite of Guava's cache and today the reference local-caching library in Java (stable 3.x line, requiring Java 11+) — uses an algorithm called W-TinyLFU that takes the best of both worlds.

The central idea of W-TinyLFU

The problem with pure LFU was that counting the frequency of everything needs a lot of memory. TinyLFU solves this: instead of an exact counter, it uses a compact probabilistic structure called a frequency sketch (a kind of Count-Min Sketch) that keeps an estimate of frequency with just a few bits per key. The "W" stands for Window: a small admission window that newcomers enter first, giving them a chance to prove themselves.

The workflow of W-TinyLFU goes like this — and it's worth understanding, because it's gold in a senior interview:

  1. Admission window: every new entry first enters a small LRU (about 1% of the cache). This gives fresh data a chance to prove itself without immediately competing with the veterans.
  2. Admission filter: when an entry wants to move from the window into the main space, TinyLFU judges like a doorman: "is this newcomer's estimated frequency higher than that of the current victim in the main space?" If yes, it's admitted and the victim goes; if no, the newcomer itself is rejected. This is what neutralizes scans — a one-shot item has low frequency and the doorman turns it away.
  3. Main SLRU space: the main space is itself a Segmented LRU with two segments: a "probation" segment and a "protected" segment. An entry that is read again is promoted from probation to protected. This combines frequency (how much) and recency (when).
  4. Counter aging: the frequency sketch periodically halves all counters (aging), so that the LFU problem of "last month's champion stubbornly staying" is solved. Past hotness slowly fades.

The result: on real workloads, Caffeine achieves a hit rate close to the ideal clairvoyant algorithm (Bélády's), while its memory overhead is negligible. You usually tune none of this by hand; you just give a maximumSize and Caffeine handles the rest intelligently.

Why you shouldn't write your own LRU

You'll be tempted to build a hand-rolled LRU with LinkedHashMap. Fine for a toy, but in production you have three problems: worse hit ratio under scans, a global lock that kills concurrency, and no TTL or stats. Caffeine solves all three and is heavily optimized — rewriting it is almost always a mistake.


Caffeine in practice: building and tuning

Let's start with a simple cache and enrich it layer by layer. The dependency (current stable from the 3.x line):

<dependency>
    <groupId>com.github.ben-manes.caffeine</groupId>
    <artifactId>caffeine</artifactId>
    <version>3.2.2</version>
</dependency>

A manual (cache-aside) cache, with a size cap and stats:

import com.github.benmanes.caffeine.cache.Cache;
import com.github.benmanes.caffeine.cache.Caffeine;
import java.time.Duration;

Cache<Long, Product> cache = Caffeine.newBuilder()
    .maximumSize(10_000)                     // size-based eviction with W-TinyLFU
    .expireAfterWrite(Duration.ofMinutes(5)) // TTL
    .recordStats()                           // count hits/misses for monitoring
    .build();

Product p = cache.getIfPresent(42L);         // may be null (miss)
cache.put(42L, product);

LoadingCache — the read-through version

If you hand the load function to Caffeine, you no longer write those three cache-aside steps yourself and — more importantly — you get automatic request coalescing for free:

import com.github.benmanes.caffeine.cache.LoadingCache;

LoadingCache<Long, Product> cache = Caffeine.newBuilder()
    .maximumSize(10_000)
    .expireAfterWrite(Duration.ofMinutes(5))
    .build(id -> productRepository.findById(id));   // load function

Product p = cache.get(42L);   // if absent, loads and stores; never null unless the loader returns null
Why LoadingCache is safer against stampedes

Suppose key 42 isn't in the cache and a thousand threads call it at once. With hand-written cache-aside, all thousand see a miss and all thousand go to the database. But LoadingCache guarantees that for a given key, only one thread runs the load function; the rest queue behind that same computation and get the ready result. This is called "request coalescing" or "single-flight," and it's your first shield against a stampede — though only within this one JVM.

TTL vs TTI: expireAfterWrite vs expireAfterAccess

Here's one of the most common interview misconceptions. There are two kinds of time-based expiry:

TTL is like milk, TTI is like a gym membership

TTL (which Caffeine calls expireAfterWrite) is like the expiry date on a milk carton: from the moment it's written, the clock starts ticking and regardless of how many times you open it, it spoils at its due time. TTI, Time To Idle (which Caffeine calls expireAfterAccess), is like a gym membership that "cancels if you don't show up for 30 days": every time you use it, the clock resets; only long idleness kills it.

Cache<String, Session> sessions = Caffeine.newBuilder()
    .expireAfterAccess(Duration.ofMinutes(30))   // TTI: stay alive while active
    .build();

Cache<String, ExchangeRate> rates = Caffeine.newBuilder()
    .expireAfterWrite(Duration.ofMinutes(1))     // TTL: an FX rate is worthless after a minute
    .build();

Rule of thumb: for data that goes stale at the source (FX rates, prices, inventory) use expireAfterWrite/TTL — no matter how often you read it, after a minute it's no longer trustworthy. For data that should be dropped only when abandoned (a user session, transient data) use expireAfterAccess/TTI.

refreshAfterWrite: the subtle difference from expiry

refreshAfterWrite differs from expireAfterWrite — this is the interview trap

expireAfterWrite is strict: after the deadline the entry is removed, and the next request must wait for a fresh load (a latent miss). But refreshAfterWrite is soft: after the deadline the stale entry is still there and is served immediately to the requester, while a fresh load kicks off in the background so it's fresh next time. So refresh doesn't make the user wait and may serve a momentary staleness; expire guarantees freshness at the cost of latency. Often you set both together: refresh shorter than expire.

LoadingCache<Long, Product> cache = Caffeine.newBuilder()
    .refreshAfterWrite(Duration.ofMinutes(1))    // after 1 min, refresh in background
    .expireAfterWrite(Duration.ofMinutes(10))    // but after 10 min, don't serve stale
    .build(id -> productRepository.findById(id));

AsyncLoadingCache and asynchronous caching

If you work in a reactive or CompletableFuture world, Caffeine also has an async version that stores the result as a CompletableFuture — which makes request coalescing work naturally over futures:

AsyncLoadingCache<Long, Product> cache = Caffeine.newBuilder()
    .maximumSize(10_000)
    .buildAsync((id, executor) ->
        CompletableFuture.supplyAsync(() -> productRepository.findById(id), executor));

CompletableFuture<Product> future = cache.get(42L);
Always take recordStats and maximumSize seriously

Two settings junior engineers forget: without maximumSize or maximumWeight, your cache is a potential memory leak that grows until an OutOfMemoryError. And without recordStats() you're blind — you don't know your hit ratio, so you don't know whether the cache even helps or is just wasting memory. An unmonitored cache is guessing, not engineering.


Local vs distributed: Caffeine vs Redis

So far we've talked about Caffeine, a local in-process cache: the data lives right inside that JVM's memory. The fastest thing possible. But it has a big problem that only shows up when you spin up multiple instances of your application.

A local cache is like each employee's pocket notebook

Imagine a company with ten employees, each with a pocket notebook (the local cache) for jotting down prices. It's fast — everyone glances at their own notebook instantly. But when a price changes, the first employee's notebook is updated and the other nine still hold the old price! Now customers get different answers depending on which employee they talk to. This "inconsistency across instances" is the biggest weakness of a local cache. The fix: one shared central ledger everyone looks at — that's Redis.

Redis is an external, shared in-memory key-value data store that lives on the network. All instances of your application connect to a single Redis, so they share one version of the truth. The cost: every access is now a network round trip (milliseconds, not nanoseconds) and data serialization (turning a Java object into bytes and back).

Feature Local cache (Caffeine) Distributed cache (Redis)
Where data lives Inside the same JVM's memory A separate server on the network
Access speed ~100 ns ~0.5–1 ms
Shared across instances No; each instance is separate Yes; one shared source
Serialization needed? No (raw objects) Yes (bytes over the wire)
What happens on app crash? Cache goes too Cache survives (it's external)
Capacity Bounded by the app's heap Much larger, independent
Extra capabilities Key-value only Data structures, pub/sub, server-side TTL, scripting
The golden rule of choosing

If you have frequently read-only data and can tolerate short-lived inconsistency, use Caffeine — it's the fastest. If you have multiple instances and need a shared, coordinated version, or the data must survive an app crash, use Redis. And in serious systems you often use both: Caffeine as the first tier (L1) and Redis as the second tier (L2) — whose architecture we'll see shortly.


Redis and its Java clients

Redis (short for REmote DIctionary Server) isn't just a key-value cache; it's an in-memory data store with rich data structures: strings, hashes, lists, sets, sorted sets, and more. But its most common role is still the cache.

Redis being "single-threaded" is a feature, not a weakness

You may hear "Redis is single-threaded." What that means is that command execution is serialized onto a single logical thread (newer versions have multi-threaded network I/O, but command processing is still serial). This is deliberate: because only one command runs at a time, Redis operations are naturally atomic and need no complex locking. That's why a command like INCR works without any race condition. Its speed comes from this same simplicity and being in-memory.

The three main Java clients you should know:

  • Lettuce: Spring Boot's default client. Built on Netty, asynchronous and reactive, thread-safe; a single connection can be shared across many threads. A good default choice.
  • Jedis: the older, simpler client — blocking and synchronous. Not thread-safe, so it's usually used with a connection pool. Simple and well-known.
  • Redisson: higher-level; instead of raw commands, it gives distributed implementations of Java structures (RMap, RLock, distributed semaphore, queue). Excellent when you need distributed locks or complex structures.

Simple usage with Lettuce (directly, without Spring):

import io.lettuce.core.RedisClient;
import io.lettuce.core.api.StatefulRedisConnection;
import io.lettuce.core.api.sync.RedisCommands;

RedisClient client = RedisClient.create("redis://localhost:6379");
try (StatefulRedisConnection<String, String> conn = client.connect()) {
    RedisCommands<String, String> cmd = conn.sync();

    cmd.set("product:42", jsonPayload);      // write
    cmd.expire("product:42", 300);           // server-side TTL: 300 seconds
    String cached = cmd.get("product:42");   // read; null means miss or expired
}
client.shutdown();
A licensing note you should know (keep it current)

Redis's licensing history has been bumpy. In 2024 Redis moved off the permissive BSD license, which led the Linux Foundation to create a fully open fork called Valkey (under BSD 3-Clause). Redis then backpedaled in version 8 and introduced a tri-license model: RSALv2, SSPLv1, and AGPLv3 (an OSI-approved license). In practice today you have two common options: Redis (version 8 onward) and Valkey (the fully-BSD fork). For a Java developer the API is nearly identical and the clients above work with either; but if you have licensing sensitivities, factor this distinction into your architecture.


Spring's cache abstraction: declarative caching with annotations

Writing that three-step pattern by hand everywhere is tedious and error-prone. Spring provides a cache abstraction that, with a few annotations, applies caching to methods declaratively and transparently — without polluting your method body.

The cache abstraction is like a memory-keeping assistant at your door

Imagine an assistant sitting at your office door. Every time someone asks a question, the assistant first checks their notebook: if they've written the answer before, they give it and don't even knock on your door (the method doesn't run). If not, they pass the question to you, take your answer, write it in the notebook, and hand it to the asker. @Cacheable is exactly that assistant — and the beauty is that you (the method) don't even know a cache is involved.

First enable the abstraction and give it a provider. With Spring Boot, add the spring-boot-starter-cache dependency and put @EnableCaching on a config class. If Caffeine is on the classpath, Spring auto-configures a CaffeineCacheManager; if Redis, a RedisCacheManager.

@Configuration
@EnableCaching
public class CacheConfig {

    @Bean
    public CacheManager cacheManager() {
        CaffeineCacheManager manager = new CaffeineCacheManager("products", "users");
        manager.setCaffeine(Caffeine.newBuilder()
            .maximumSize(10_000)
            .expireAfterWrite(Duration.ofMinutes(5))
            .recordStats());
        return manager;
    }
}

@Cacheable — cached reads

@Service
public class ProductService {

    @Cacheable(cacheNames = "products", key = "#id")
    public Product getProduct(long id) {
        // this body runs only on a miss
        return productRepository.findById(id);
    }
}

The first call with id = 42 runs the body and caches the result; subsequent calls with the same key don't run the body at all and return straight from the cache.

The classic trap: self-invocation via the proxy

Spring's cache abstraction works via a proxy: Spring wraps your bean in a layer that checks the cache. If a method in the same class calls another @Cacheable method with this.getProduct(...), the call doesn't go through the proxy and the cache is completely bypassed — the body runs every time as if there were no cache. This is one of the most common silent bugs. The fix: make the call from a different bean, or move that method to a separate service.

@CacheEvict and @CachePut — invalidation and update

  • @CacheEvict removes an entry from the cache — for when data changed or was deleted.
  • @CachePut always runs the body and puts its result in the cache (declarative write-through).
@CachePut(cacheNames = "products", key = "#product.id")
public Product update(Product product) {
    return productRepository.save(product);   // both DB and cache are updated
}

@CacheEvict(cacheNames = "products", key = "#id")
public void delete(long id) {
    productRepository.deleteById(id);
}

@CacheEvict(cacheNames = "products", allEntries = true)
public void reloadAll() {
    // clears the entire products cache
}

Other annotations: @Caching to combine multiple operations on one method, and @CacheConfig to put shared settings (like a cache name) at the class level. Also condition (cache only if a condition holds) and unless (do not cache if a condition holds, evaluated against the result) give fine control:

@Cacheable(cacheNames = "products", key = "#id",
           unless = "#result == null")   // don't cache a null result
public Product getProduct(long id) { ... }
The sync flag — Spring's built-in shield against stampedes

By default @Cacheable does not protect against concurrent misses: if a thousand threads want a cold key at once, all thousand run the body. With @Cacheable(sync = true) Spring guarantees that for one key only one thread runs the body while the rest wait for its result — that's request coalescing. It's the simplest in-JVM defense against a cache stampede. Remember it only works within a single instance, not across instances.


Two-tier architecture: L1 (Caffeine) + L2 (Redis)

In serious systems you often want both worlds: local nanosecond speed and cross-instance sharing. This becomes a multi-tier cache or near cache.

L1 is your pocket, L2 is your bag, the database is home

You keep money in three places. Your pocket (L1 / Caffeine) has the fastest access but little room. Your bag (L2 / Redis) is slower but roomier and shared with others. And home (the database) is the ultimate source. When you want something: pocket first, then bag, then home — and every time you fetch from the depths, you also stash it in the higher layers so next time is faster. That's exactly the read flow of a two-tier cache.

public Product get(long id) {
    // L1: ultra-fast local cache
    Product p = caffeine.getIfPresent(id);
    if (p != null) return p;

    // L2: shared Redis cache
    p = readFromRedis(id);
    if (p != null) {
        caffeine.put(id, p);   // populate L1 too
        return p;
    }

    // primary source
    p = productRepository.findById(id);
    if (p != null) {
        writeToRedis(id, p);   // fill L2
        caffeine.put(id, p);   // and L1
    }
    return p;
}

But two tiers bring a serious problem: L1 invalidation. When data changes, Redis (shared) can be cleared directly, but each instance has its own independent Caffeine that doesn't know. The common fix: use Redis pub/sub — when data changes, publish a message on a Redis channel; all instances subscribe to that channel and, on receiving the message, evict the corresponding entry from their own local Caffeine.

Near cache in ready-made libraries

You don't always have to build this by hand. Redisson provides an RLocalCachedMap that implements exactly this near cache with automatic pub/sub-based invalidation. But understanding the mechanism (local L1 + shared L2 + an invalidation channel) matters more than knowing the class name — because in an interview it's the mechanism they'll ask about.


Cache stampede: the thundering herd and how to tame it

Now we reach one of the most dangerous cache phenomena, which we've alluded to many times. It has many names: cache stampede, thundering herd, dogpile.

A cache stampede is like everyone rushing one door at once

Imagine a very hot key — say, the storefront's front-page price, read a thousand times a second — suddenly hits its TTL and leaves the cache. At that instant, a thousand requests that wanted that key simultaneously see a miss and simultaneously storm the database to rebuild it. A database that comfortably handles a thousand cached requests now buckles under a thousand heavy concurrent queries — sometimes dragging the whole system down with it. The worst part? All thousand are doing exactly the same work.

Several proven ways to tame it:

1) Request coalescing / single-flight. Exactly what LoadingCache and @Cacheable(sync=true) give: guarantee that for one key only one computation runs while the rest queue behind it. The simplest, first line of defense — but only within one JVM.

2) Distributed lock (mutex). For a cross-instance stampede, allow only one instance in the whole cluster to rebuild the key; the rest either wait briefly or serve the previous stale value. With Redis you build this via SET key value NX PX <ms> — an atomic, timed lock:

// only one instance acquires the lock and rebuilds
boolean gotLock = "OK".equals(
    cmd.set("lock:product:42", token, SetArgs.Builder.nx().px(3000)));
if (gotLock) {
    try {
        Product fresh = productRepository.findById(42L);
        writeToRedis(42L, fresh);
    } finally {
        // safe release (only if you still own it) via a Lua script
        releaseLockIfOwner("lock:product:42", token);
    }
}

3) Probabilistic early expiration (XFetch). A subtle, beautiful technique: instead of everyone expiring at one exact instant, the closer you get to expiry, the higher the probability each request decides to rebuild the cache early (in the background). Because this decision is independent and random, usually only one of the thousand requests rebuilds early while the rest still get a valid cache. The academic formula: rebuild if now - delta * beta * ln(random()) >= expiry, where delta is the recompute time and beta (default 1.0) is a tuning parameter.

4) TTL jitter. If you fill a thousand keys together with the exact same TTL, they all expire simultaneously and create a mass stampede. The simple fix: add a bit of randomness to each TTL (e.g. 5 minutes ± 30 seconds) so expiries are spread out.

5) Stale-while-revalidate. Keep the stale data and serve it immediately, refreshing in the background. That's exactly what refreshAfterWrite does in Caffeine.

Which one to choose?

For a stampede within one JVM, LoadingCache or sync=true is usually enough — simple and free. For a cross-instance stampede on a hot, expensive key, you need a Redis distributed lock or XFetch. And apply jitter almost always; it's cheap and prevents mass stampedes caused by synchronized expiry. These aren't mutually exclusive; you often combine several.

Don't confuse the three cache demons

There are three different problems with similar-sounding names. Cache stampede (as described): a simultaneous rush after a hot key expires. Cache penetration: repeated requests for a key that doesn't exist at all (e.g. id = -1), so every one misses and goes straight to the database — fix: cache the null/empty result too with a short TTL, or add a Bloom filter. Cache avalanche: when a large portion of the cache expires at once, or Redis itself crashes, and suddenly all the load pours onto the database — fix: jitter, Redis redundancy, and a circuit breaker. In an interview, distinguishing these three is a big win.


Invalidation and the consistency trade-off

We arrive at that "hardest problem." When data changes at the source, how does the cache know its copy is stale? A few strategies:

  • TTL-based (expiring): the simplest. Every entry dies on its own after a while. Consistency is "eventual": the staleness window is at most the TTL. For most data this is enough and surprisingly effective.
  • Event-based (explicit): when data changes, actively evict or update the cache entry (@CacheEvict/@CachePut). More precise but more fragile: you must find every data-change path and miss none — hard in practice.
  • Versioned key: instead of deleting, change the key. For example product:42:v7; when the product changes, bump the version to v8. Now readers looking for v8 naturally miss, and the old key expires unused. Elegant and free of deletion coordination.
Why invalidation is so hard

The deep problem with explicit invalidation is that it's a distributed consistency problem. Between "writing to the database" and "clearing the cache" there's a time gap, and in that gap — or if the clear fails — the cache serves wrong data. Even ordering matters: if you clear the cache first and then write the database, a concurrent reader can put the old database value back into the cache and leave it stale forever. That's why many seasoned systems, rather than chasing perfect consistency, rely on a short TTL: instead of "always correct," they accept "at most N seconds wrong," because it's far simpler to control.

The consistency trade-off is like a newspaper vs an airport's live board

You have two mental models for freshness. A newspaper (a TTL cache): it's a snapshot of the moment it was printed; you know it may get a bit stale by evening, but it's cheap, quickly available, and good enough for most things. An airport's live flight board (no cache, or write-through): always instant and precise, but very expensive to maintain. The right engineering question isn't "which is better," but "how much staleness is acceptable for this data?" A stock price? Seconds. A product category name? An hour is fine.

The final architectural note: caching almost always means accepting eventual consistency. The moment you keep a copy, you create a window of divergence. A good engineer's job isn't to eliminate this window (often impossible or very expensive), but to deliberately shrink it to a size that's harmless for the problem domain.


Pitfalls and best practices

Let's gather the top traps in one place — the ones that bite in real code or interviews:

  • An unbounded cache = a memory leak. Always set maximumSize/maximumWeight (or a TTL); an unbounded cache will eventually throw OutOfMemoryError.
  • Self-invocation bypasses the cache. Calling this.method() inside the same bean skips Spring's caching layer.
  • Caching null carelessly. If you cache null uncontrolled and the source later gets a value, it stays stale; if you never cache it, you're vulnerable to penetration. Make a deliberate decision (often: cache null with a short TTL).
  • Same TTL for everything = avalanche. Add jitter so expiries spread out.
  • Caching a mutable object and sharing it. If the object you cached is mutable and the caller mutates it, the cached version changes too (in a local cache). Cache immutable objects or hand out copies.
  • Fragile serialized keys in Redis. If you store a Java object via Java native serialization in Redis, a class change can break deserialization and — worse — open an attack surface. Use JSON or an explicitly versioned format.
  • Forgetting monitoring. Without a hit ratio you don't know your cache works. Take recordStats and metrics seriously.

Best practices:

  • For every cache, answer these three questions: What's the size cap? What's the expiry policy (TTL or TTI, and how long)? How do I see my hit ratio?
  • Local cache for frequently read-only data; Redis for cross-instance sharing; two-tier for both.
  • For stampedes, start with request coalescing (sync/LoadingCache), and add a distributed lock or XFetch for hot cross-instance keys.
  • Rather than chasing perfect consistency, accept a short TTL except where staleness is truly intolerable.

Interview Questions

Now let's drill everything through real senior-interview questions. Answer each yourself first, then open the answer.

1) What exactly is the difference between cache-aside and read-through?

In cache-aside the application code is responsible for everything: it reads the cache, on a miss reads the source, and itself puts the result in the cache. The cache is just a passive key-value store. In read-through that logic is delegated to the cache: you provide a load function and the cache calls it on a miss and fills itself. The important practical difference: read-through usually coalesces concurrent requests for the same key, so it's safer against stampedes; cache-aside has no automatic protection. Caffeine's LoadingCache is a read-through example.

2) What's the difference between TTL and TTI, and which Caffeine method is each?

TTL (Time To Live) means an entry stays alive at most for a fixed time from the moment it's written, regardless of access count — in Caffeine that's expireAfterWrite. Suited to data that goes stale at the source (FX rates, prices). TTI (Time To Idle) means an entry stays alive as long as it's used and only long idleness kills it; each access resets the clock — in Caffeine that's expireAfterAccess. Suited to user sessions or transient data. You can set both together.

3) (Tricky) What's the difference between expireAfterWrite and refreshAfterWrite?

expireAfterWrite is hard: after the deadline the entry is removed and the next request waits for a fresh load (a latent miss on the critical path). refreshAfterWrite is soft: after the deadline the stale value is still there and returned immediately, while a fresh load runs in the background to be up to date next time. So refresh eliminates user latency at the cost of momentary staleness; expire guarantees freshness at the cost of latency. Often you set both together, with refresh shorter than expire.

4) What is W-TinyLFU and why is it better than LRU and LFU?

W-TinyLFU is Caffeine's eviction algorithm that combines frequency and recency. The problem with LRU is that it collapses under a sequential scan (one-shot data pushes out the hot data). The problem with LFU is heavy memory for counters and getting stuck on past hotness. W-TinyLFU solves both: a frequency sketch (like a Count-Min Sketch) estimates frequency with a few bits; an admission window gives newcomers a chance; an admission filter only admits an entry whose estimated frequency beats the victim's (neutralizing scans); and counter aging forgets stale hotness. The result: a hit rate close to optimal with negligible memory overhead.

5) When do you use a local cache (Caffeine) and when a distributed one (Redis)?

Caffeine when data is frequently read-only, you tolerate short-lived cross-instance inconsistency, and you want maximum speed (nanoseconds, no network or serialization). Redis when multiple instances must see a shared, coordinated version, or the cache must survive an app crash, or the data volume exceeds the heap. In serious systems, often both together: Caffeine as L1 and Redis as L2 (near cache). Golden line: "local for speed, distributed for coordination."

6) What is a cache stampede (thundering herd) and how do you tame it?

When a hot key expires and thousands of concurrent requests simultaneously miss and simultaneously storm the database to rebuild it, the database buckles under the redundant work. Ways to tame: (1) Request coalescing/single-flight with LoadingCache or @Cacheable(sync=true) — one computation per key, within one JVM. (2) Distributed lock with Redis SET NX PX for a cross-instance stampede. (3) Probabilistic early expiration (XFetch) so randomly only one rebuilds early. (4) TTL jitter to spread out expiries. (5) Stale-while-revalidate (like refreshAfterWrite). You usually combine several.

7) (Distinguish) Separate stampede, penetration, and avalanche.

Stampede/thundering herd: a simultaneous rush after one hot key expires. Fix: coalescing, lock, XFetch. Penetration: repeated requests for a key that doesn't exist at all, so every one misses and goes straight to the database. Fix: cache the null result with a short TTL, or a Bloom filter. Avalanche: a large portion of the cache expiring at once (or Redis itself crashing), dumping all the load onto the database at once. Fix: jitter, Redis redundancy, circuit breaker. Precisely distinguishing these three signals seniority.

8) Why does @Cacheable sometimes "do nothing" and the cache gets bypassed?

Almost always because of self-invocation. Spring's cache abstraction works via a proxy; the cache layer only kicks in when the call comes from outside the bean and through the proxy. If a method in the same class calls another @Cacheable method with this.method(...), it doesn't go through the proxy and the cache is completely bypassed — the body runs every time. Fix: make the call from another bean or move the method to a separate service. Other likely causes: missing @EnableCaching, or a condition/unless that prevented caching.

9) What are the invalidation strategies and which do you prefer?

Three main strategies: TTL (each entry dies on its own; simple and robust, but staleness up to the TTL), event-based (explicitly evict/update with @CacheEvict/@CachePut; precise but fragile since you must cover every change path), and versioned key (instead of deleting, bump the key's version so readers naturally miss). In practice I often prefer a short TTL because explicit invalidation is a distributed consistency problem (the gap between the DB write and the cache clear, and the risk the clear fails); you can also combine TTL and events.

10) (Tricky ordering) If you want to write the database and clear the cache, which first?

Write the database first, then clear the cache (cache-aside invalidation). If you do it the other way — clear the cache first, then write the DB — a concurrent reader can, between the two operations, read the old database value and put it back into the cache, leaving it stale forever. Even "write DB then clear cache" isn't flawless (if the clear fails the cache stays stale), which is why we usually keep a TTL as a safety net. This is also why "invalidate rather than update" is popular: clearing is simpler than keeping in sync.

11) Why is Redis called "single-threaded" and what does that imply?

It means command processing is serialized onto a single logical thread (newer versions have multi-threaded network I/O, but command execution itself is still serial). Positive implication: each command is naturally atomic without complex locking, so a command like INCR or SET NX works with no race condition — which is what makes a distributed lock possible. Negative implication: one heavy command (like KEYS * on a large database) blocks the whole server, so you must avoid slow commands. Overall speed comes from being in-memory plus this simplicity.

12) (Memory trap) Why is caching a mutable object in Caffeine dangerous?

Because a local cache holds the object's reference, not a copy. If you hand out the cached object to a caller and they mutate one of its fields, the version inside the cache changes too — since both point to the same object. Result: the cache serves tampered, unintended data and creates deeply confusing bugs. Fix: cache immutable objects (like a record with defensive copies), or copy on read/write. In Redis this isn't an issue because data is serialized and copied, but there you pay the serialization cost.

13) (Architecture) How does a two-tier cache work and what's its biggest challenge?

L1 is a fast local cache (Caffeine) inside each instance, and L2 is a shared cache (Redis). Reads: L1 first, then L2, then the database; and every time you fetch from the depths you also fill the higher layers. Benefit: local speed + cross-instance sharing. The biggest challenge is L1 invalidation: when data changes, L2 (shared) can be cleared directly, but each instance has its own independent L1 that doesn't know. The common fix: Redis pub/sub — on a change, publish a message and have all instances evict the corresponding entry from their own L1. Libraries like Redisson (RLocalCachedMap) provide this ready-made.

14) What hidden costs does caching have, and when should you *not* cache?

Hidden costs: (1) Inconsistency — the moment you keep a copy you create a staleness window. (2) Operational complexity — another stateful system to monitor, measure, and debug. (3) Memory usage and leak risk if unbounded. (4) Serialization cost in a distributed cache. You should not cache when: data changes fast and staleness is unacceptable; the hit ratio is low (a cold cache just wastes memory and complexity); the primary source is already fast and cheap enough; or the data is sensitive and keeping a copy is a security risk. A cache is a tool, not an obligation — measure the hit ratio first, then judge.


In a nutshell
  • Why cache: the fastest work is the work not done. A cache reduces both latency and primary-source load — often the real goal is protecting the database from collapse.
  • Patterns: cache-aside (manual, simple), read-through (cache with a load function and auto-coalescing), write-through (fresh but slow writes), write-behind (fast but data-loss risk), refresh-ahead (proactive refresh).
  • Eviction: LRU breaks under scans, LFU gets stuck on the past; W-TinyLFU in Caffeine takes the best of both with a frequency sketch + window + admission filter + aging.
  • Caffeine: a nanosecond local cache; expireAfterWrite (TTL) vs expireAfterAccess (TTI), and refreshAfterWrite which serves stale immediately and refreshes in the background. Always set maximumSize and recordStats.
  • Redis: an external shared cache for cross-instance coordination and durability beyond a crash; clients: Lettuce (default), Jedis, Redisson. Serial command execution ⇒ natural atomicity.
  • Spring: @Cacheable/@CacheEvict/@CachePut give declarative caching; beware self-invocation; sync=true is the in-JVM shield against stampedes.
  • Stampede and friends: stampede (coalescing/lock/XFetch), penetration (cache null/Bloom filter), avalanche (jitter/redundancy). Apply jitter almost always.
  • Consistency: caching means accepting eventual consistency. Often a short TTL beats chasing fragile perfect consistency. Write the DB first, then clear the cache.