Concurrency · همزمانی سنیورSenior ~69 دقیقه مطالعه~58 min read

دام‌ها و الگوهای همزمانیConcurrency Pitfalls & Patterns

در این درس یاد می‌گیری چرا برنامه‌های چندنخی خراب می‌شوند — بن‌بست، زنده‌قفلی، گرسنگی و رقابت — و با تشبیه و کد جاوای معیوب-سپس-اصلاح‌شده، الگوهای آزموده‌ای را می‌آموزی که هرکدام را خنثی می‌کنند.In this lesson you'll learn why multithreaded programs break — deadlock, livelock, starvation, and races — and, through analogies and buggy-then-fixed Java, the battle-tested patterns that defeat each one.

پیش‌نیاز:Prerequisites: همگام‌سازی، قفل‌ها، AQS و مدل حافظهٔ جاواSynchronization, Locks, AQS & the Java Memory Model


خب، بیا با هم روراست باشیم: همزمانی (concurrency) جایی است که برنامه‌نویس‌های خوب هم زانو می‌زنند. کدی که تنها روی یک نخ (thread) بی‌عیب کار می‌کند، به‌محض اینکه دو نخ همزمان به یک داده دست می‌زنند، می‌تواند پاسخ غلط بدهد یا برای همیشه معلق شود. خبر خوب این است که تمام این فاجعه‌ها در چند خانوادهٔ محدود جا می‌گیرند، و برای هر خانواده یک الگوی آزموده وجود دارد. در این درس قرار است این نقشه را کامل در ذهنت بسازیم — از صفر.

نقشهٔ راه این درس

اول یک قطب‌نمای بزرگ می‌سازیم: هر باگ همزمانی یا ایمنی (safety) را می‌شکند یا زندگی‌مندی (liveness) را. بعد سراغ چهار هیولا می‌رویم: بن‌بست (deadlock)، زنده‌قفلی (livelock)، گرسنگی (starvation) و شرایط رقابتی (race condition). سپس ابزارهای دفاعی را یکی‌یکی یاد می‌گیریم: صف مسدودکننده (BlockingQueue) برای تولیدکننده-مصرف‌کننده، اولیه‌های هماهنگی (latch/barrier/semaphore/phaser)، و در آخر دو راهبرد ساختاری که کلاس‌های کاملی از باگ را از ریشه حذف می‌کنند — تغییرناپذیری (immutability) و محصورسازی (confinement). درس با یک بخش کامل سؤالات مصاحبه بسته می‌شود.

بخش ۰ — واژه‌هایی که باید بلد باشی

قبل از هر چیز، چند واژه را که مدام تکرار می‌شوند، همین‌جا با تشبیه باز کنیم تا بعداً سردرگم نشوی.

  • نخ (thread): یک خط اجرای مستقل. تصور کن هر نخ یک کارگر است که همزمان با کارگرهای دیگر مشغول است.
  • قفل (lock): کلید یک اتاق. هر لحظه فقط یک کارگر می‌تواند کلید را داشته باشد؛ بقیه پشت در منتظر می‌مانند. در جاوا واژهٔ synchronized و کلاس ReentrantLock همین کلید را می‌سازند.
  • بازورودپذیر (reentrant): کارگری که کلید اتاق را دارد، می‌تواند بدون قفل‌شدن دوباره وارد همان اتاق شود. قفل‌های synchronized جاوا بازورودپذیرند.
  • اتمی (atomic): عملیاتی که یا کامل انجام می‌شود یا اصلاً — هیچ کارگر دیگری نمی‌تواند وسط کار، حالتِ نیمه‌تمام را ببیند.
  • درهم‌آمیزی (interleaving): ترتیبی که سیستم‌عامل قدم‌های کارگرهای مختلف را در هم می‌بافد. تو هیچ کنترلی روی این ترتیب نداری، و درست همین‌جاست که باگ‌ها زاده می‌شوند.
آشپزخانهٔ یک رستوران

تمام این درس را می‌توانی با تصویر یک آشپزخانهٔ شلوغ بفهمی: نخ‌ها آشپزها هستند، قفل‌ها ابزارهای مشترک (یک چاقو، یک اجاق)، و داده‌های مشترک همان مواد غذایی روی میز. وقتی دو آشپز همزمان یک قابلمه را بدون هماهنگی بردارند، یا غذا خراب می‌شود (نقض ایمنی) یا هر دو منتظر هم می‌مانند و هیچ غذایی آماده نمی‌شود (نقض زندگی‌مندی).

مدل ذهنی: زندگی‌مندی (liveness) در برابر ایمنی (safety)

بیا با بزرگ‌ترین ایده شروع کنیم، چون بقیهٔ درس زیر سایهٔ آن است. هر باگ همزمانی، بدون استثنا، نقض یکی از این دو ویژگی است:

  • ایمنی (safety) یعنی «هیچ اتفاق بدی هرگز رخ نمی‌دهد.» یک شرط رقابتی (race condition) وضعیت را خراب می‌کند، یک HashMap هنگام تغییر اندازه (resize) وارد حلقهٔ بی‌نهایت می‌شود، یک الگوی «بررسی-سپس-عمل» (check-then-act) مقدار کهنه می‌بیند. شکست ایمنی پاسخ غلط تولید می‌کند.
  • زندگی‌مندی (liveness) یعنی «سرانجام اتفاق خوبی می‌افتد.» بن‌بست، زنده‌قفلی و گرسنگی یعنی نخ‌ها دیگر پیشرفت نمی‌کنند. شکست زندگی‌مندی معلق‌شدن (hang) تولید می‌کند.

چرا این تفکیک این‌قدر مهم است؟ چون راه‌حل‌هایشان روحیهٔ متضاد دارند و اگر این را نفهمی، رفعِ یک باگ باگِ دیگری می‌سازد.

دو نوع خرابی در آشپزخانه

شکست ایمنی مثل این است که دو آشپز همزمان نمک بریزند و غذا شور از آب دربیاید — غذا آماده شد، ولی غلط است. شکست زندگی‌مندی مثل این است که دو آشپز هر دو منتظر بمانند تا دیگری اول اجاق را رها کند — غذا هرگز آماده نمی‌شود، هرچند هیچ‌کدام خطای آشکاری نکرده‌اند. یکی «نتیجهٔ غلط» است، دیگری «هیچ نتیجه‌ای».

مهندسان ارشد این تفکیک را درونی می‌کنند چون راه‌حل‌ها متضادند. ایمنی معمولاً به هماهنگی بیشتر نیاز دارد (قفل، atomic، رابطهٔ happens-before). زندگی‌مندی معمولاً به هماهنگی کمتر یا هوشمندتر نیاز دارد (ترتیب‌دهی قفل، مهلت‌زمانی، انصاف، الگوریتم‌های بدون‌قفل). زیاده‌روی در همگام‌سازی برای رفع یک رقابت می‌تواند بن‌بست بسازد؛ شل‌کردن قفل‌ها برای رفع بن‌بست می‌تواند رقابت را برگرداند.

تو همیشه روی یک لبهٔ باریک تعادل برقرار می‌کنی

از یک سو رقابت (نیاز به قفلِ بیشتر)، از سوی دیگر بن‌بست (نیاز به قفلِ کمتر یا هوشمندتر). هدف مهندسی خوب، نه پرتاب‌شدن به هیچ‌کدام از دو طرف، بلکه ماندن روی همان لبه است. هر وقت خواستی یک قفل اضافه کنی، از خودت بپرس: «آیا این یک بن‌بست جدید نمی‌سازد؟»

زیربنای نامرئی: مدل حافظهٔ جاوا (JMM)

حالا یک لایهٔ عمیق‌تر. مدل حافظهٔ جاوا (Java Memory Model یا JMM) زیربنای همهٔ این‌هاست. اصطلاح کلیدی‌اش happens-before است: یک «یال» یا رابطهٔ تضمینی بین دو عمل، که می‌گوید عمل اول قطعاً پیش از عمل دوم دیده می‌شود.

یادداشت روی تخته‌سفید

تصور کن کارگر A چیزی روی دفترچهٔ شخصی‌اش می‌نویسد. کارگر B تا وقتی A آن را روی تخته‌سفید مشترک منتشر نکند، ممکن است نسخهٔ قدیمی یا نیمه‌نوشته را ببیند — یا اصلاً نبیند. رابطهٔ happens-before همان «انتشار روی تخته‌سفید» است: تضمینی که آنچه A نوشت، به‌درستی و کامل به چشم B می‌رسد.

بدون یک یال happens-before بین یک نوشتن روی نخ A و یک خواندن روی نخ B، ممکن است B مقدار کهنه، شیء نیمه‌ساخته، یا عملیات بازچینش‌شده (reordered) ببیند — حتی روی معماری x86. چه چیزهایی این یال را برقرار می‌کنند؟ قفل‌ها (synchronized، ReentrantLock)، کلیدواژهٔ volatile، فیلدهای final (پس از پایان ساخت)، Thread.start/join، و کلاس‌های java.util.concurrent.

«روی ماشین من کار می‌کرد» یک استدلال نیست

اگر نتوانی توضیح دهی که کدام رابطهٔ happens-before تضمین می‌کند نخ B نوشتهٔ نخ A را ببیند، کد تو صرفاً بخت آورده — و بخت روی سخت‌افزار دیگر یا زیر بار سنگین برمی‌گردد. همیشه استدلال happens-before داشته باش.


بن‌بست (deadlock): چهار شرط کافمن (Coffman)

چهارراهِ بدون چراغ

تصور کن چهار ماشین از چهار جهت به یک چهارراه بدون چراغ می‌رسند و هرکدام می‌خواهد وارد شود اما منتظر است تا ماشین سمت راستش اول برود. هیچ‌کدام حرکت نمی‌کند چون هرکدام منتظر دیگری است. این دقیقاً بن‌بست است: یک چرخهٔ انتظار که هرگز باز نمی‌شود.

بن‌بست یک چرخه از نخ‌هاست که هرکدام منبعی را در دست دارند که نخ بعدی می‌خواهد. نکتهٔ فوق‌العادهٔ کاربردی این است که برای وقوع بن‌بست، به هر چهار شرط کافمن به‌طور همزمان نیاز است — و اگر فقط یکی را بشکنی، بن‌بست به‌کلی ناممکن می‌شود.

شرط معنا چگونه بشکنیمش
انحصار متقابل (mutual exclusion) منبع به‌صورت انحصاری در دست است منابع تغییرناپذیر/اشتراکی-خواندنی، ساختارهای بدون‌قفل
نگه‌داشتن و انتظار (hold and wait) نخ یک قفل را نگه می‌دارد و قفل دیگری می‌خواهد همهٔ قفل‌ها را یکجا بگیر، یا پیش از درخواست آزاد کن
عدم پیش‌دستی (no preemption) قفل‌ها را نمی‌توان به‌زور گرفت از tryLock با مهلت + عقب‌نشینی استفاده کن
انتظار دوری (circular wait) چرخه‌ای در گراف انتظار وجود دارد یک ترتیب سراسری قفل‌ها تحمیل کن

بیا کلاسیک‌ترین باگ بن‌بست را ببینیم: دو حساب بانکی و دو نخ که در جهت‌های مخالف پول انتقال می‌دهند.

// باگ: ترتیب قفل به ترتیب آرگومان وابسته است → انتظار دوری
void transfer(Account from, Account to, long amount) {
    synchronized (from) {
        synchronized (to) {          // T1: A سپس B؛ T2: B سپس A → بن‌بست
            from.debit(amount);
            to.credit(amount);
        }
    }
}

چه اتفاقی می‌افتد؟ نخ اول transfer(A, B, ...) را صدا می‌زند و قفل A را می‌گیرد. دقیقاً همان لحظه نخ دوم transfer(B, A, ...) را صدا می‌زند و قفل B را می‌گیرد. حالا نخ اول منتظر B است (که دست نخ دوم است) و نخ دوم منتظر A است (که دست نخ اول است). چرخهٔ انتظار کامل شد؛ هر دو تا ابد مسدود.

ترتیب آرگومان‌ها تصمیم‌گیرندهٔ ترتیب قفل شده — و این فاجعه است

ریشهٔ باگ این است که ترتیب قفل‌گیری به ترتیبی که فراخواننده آرگومان‌ها را داده گره خورده. تا وقتی دو فراخوانی با ترتیب آرگومانِ معکوس وجود داشته باشد، انتظار دوری کمین کرده. راه‌حل باید این وابستگی را قطع کند.

راه‌حل ۱: ترتیب سراسری قفل‌ها (شرط انتظار دوری را می‌شکند)

ایده ساده و زیباست: به هر شیء قابل‌قفل یک کلید ترتیب یکتا و پایدار بده و همیشه به همان ترتیب قفل بگیر، فارغ از اینکه فراخواننده چه ترتیبی داده.

void transfer(Account from, Account to, long amount) {
    Account first  = from.id() < to.id() ? from : to;  // ترتیب کلی بر اساس id
    Account second = from.id() < to.id() ? to : from;
    synchronized (first) {
        synchronized (second) {
            from.debit(amount);
            to.credit(amount);
        }
    }
}
چرا این جواب می‌دهد

اگر همه همیشه اول قفلِ با id کوچک‌تر را بگیرند، دیگر ممکن نیست دو نخ در جهت مخالف قفل بگیرند. چرخه‌ای که برای بن‌بست لازم است هرگز شکل نمی‌گیرد. این «ترتیب سراسری» عملی‌ترین سلاح ضدبن‌بست در بیشتر سیستم‌های واقعی است.

حالا یک حالت مرزی: اگر from.id() == to.id() (انتقال به‌خود) باشد، همان مانیتور را دوبار synchronized می‌کنی — بی‌ضرر است چون مانیتورهای جاوا بازورودپذیر (reentrant) هستند (یادت هست؟ کارگری که کلید را دارد می‌تواند دوباره وارد شود)، اما باید معنایی هم از انتقال به‌خود محافظت کنی. و وقتی هیچ کلید یکتای طبیعی وجود ندارد، از System.identityHashCode و یک قفل شکنندهٔ تساوی (tie-breaker) برای برخورد (collision) نادرِ هش‌های برابر استفاده کن:

private static final Object TIE = new Object();

void transfer(Account from, Account to, long amount) {
    int hf = System.identityHashCode(from), ht = System.identityHashCode(to);
    if (hf < ht)       lockedTransfer(from, to, amount);
    else if (hf > ht)  lockedTransfer(to, from, amount, /*reverse*/ true);
    else synchronized (TIE) { lockedTransfer(from, to, amount); } // برخورد هش
}

قفل TIE آن حالت نادر را می‌گیرد که دو شیء متفاوت اتفاقاً identityHashCode برابر داشته باشند و ترتیبشان قابل تعیین نباشد؛ در آن یک مورد نادر، یک قفل مشترک واحد تضمین می‌کند ترتیب همچنان یکنواخت بماند.

راه‌حل ۲: tryLock با مهلت (شرط عدم پیش‌دستی را می‌شکند)

راهبرد دوم به‌جای مرتب‌کردن قفل‌ها، به نخ اجازه می‌دهد از انتظار جاودانه فرار کند: به‌جای «تا هر وقت شد صبر کن»، بگو «حداکثر ۵۰ میلی‌ثانیه تلاش کن، نشد بی‌خیال شو و دوباره امتحان کن».

boolean transfer(Account from, Account to, long amount, Duration timeout)
        throws InterruptedException {
    long deadline = System.nanoTime() + timeout.toNanos();
    while (System.nanoTime() < deadline) {
        if (from.lock.tryLock(50, TimeUnit.MILLISECONDS)) {
            try {
                if (to.lock.tryLock(50, TimeUnit.MILLISECONDS)) {
                    try { from.debit(amount); to.credit(amount); return true; }
                    finally { to.lock.unlock(); }
                }
            } finally { from.lock.unlock(); } // همیشه قفل بیرونی را آزاد کن
        }
        // هر دو را نگرفتیم: مقداری تصادفی عقب بکش تا از زنده‌قفلی جلوگیری شود
        Thread.sleep(ThreadLocalRandom.current().nextInt(1, 10));
    }
    return false;
}
جزئیاتِ حیاتی، همان `finally` بیرونی است

اگر to را نتوانی بگیری اما قفل from را در دست نگه داری و دوباره تلاش کنی، تازه شرط «نگه‌داشتن و انتظار» را ساخته‌ای — یعنی دقیقاً همان بن‌بستی که می‌خواستی از آن فرار کنی! آن finally که from را آزاد می‌کند، یک شکست گذرا را به یک عقب‌نشینی تمیز تبدیل می‌کند، نه یک قفل ابدی. و عقب‌نشینی تصادفی (نه ثابت) همان چیزی است که مانع می‌شود حلقهٔ تلاش‌مجدد به زنده‌قفلی تبدیل شود.

تشخیص بن‌بست در محیط تولید (production)

فرض کن با وجود همهٔ احتیاط‌ها، سرور در محیط واقعی معلق شده. چطور بفهمی بن‌بست است؟

  • Thread dump: دستور jstack <pid> (یا kill -3) بخش «Found one Java-level deadlock» را با چرخهٔ دقیق چاپ می‌کند. این نخستین اقدام تو روی هر JVM معلق است.
  • برنامه‌نویسانه: ThreadMXBean.findDeadlockedThreads() می‌تواند روی یک نخ نگهبان (watchdog) اجرا شود و خودکار هشدار دهد.
ThreadMXBean mx = ManagementFactory.getThreadMXBean();
long[] deadlocked = mx.findDeadlockedThreads(); // اگر نبود null
if (deadlocked != null) log.error("DEADLOCK: {}", Arrays.toString(deadlocked));

زنده‌قفلی (livelock) و گرسنگی (starvation)

بن‌بست تنها راه معلق‌شدن نیست. دو خویشاوند نزدیک دارد که فریبنده‌ترند چون نخ‌ها به‌ظاهر مشغول‌اند.

دو نفر در راهرو

دو نفر در یک راهروی باریک روبه‌روی هم قرار می‌گیرند. هر دو مؤدبانه به یک سمت کنار می‌روند — باز روبه‌روی هم. دوباره هر دو به سمت دیگر — باز روبه‌روی هم. آن‌ها مسدود نیستند، دارند فعالانه حرکت می‌کنند، اما هرگز رد نمی‌شوند. این زنده‌قفلی است.

زنده‌قفلی (livelock): نخ‌ها مسدود نیستند — فعالانه در حال اجرا هستند — اما مدام به یکدیگر واکنش نشان می‌دهند و هیچ پیشرفتی ندارند. در کد وقتی ظاهر می‌شود که نخ‌ها هم‌گام عقب می‌کشند و تلاش دوباره می‌کنند، یا بازیگرهای (actor) پیام‌رسان یک وظیفه را مدام به هم رد می‌کنند. درمانش عدم تقارن (asymmetry) است: عقب‌نشینی تصادفی (همان nextInt(1, 10) بالا)، یا یک اولویت/توکن که تقارن را می‌شکند. اگر هر دو نفر در راهرو یک سکه بیندازند تا تصمیم بگیرند چه کسی اول برود، تقارن شکسته و مسئله حل می‌شود.

صف نانوایی بی‌نظم

گرسنگی مثل نانوایی‌ای است که صف مرتب ندارد و هر بار هرکس زورش بیشتر است نان می‌گیرد. یک آدم مؤدب و آرام ممکن است ساعت‌ها بایستد و هرگز نوبتش نشود — نه چون قفل شده، بلکه چون دیگران همیشه او را کنار می‌زنند.

گرسنگی (starvation): یک نخ هرگز منبعی را نمی‌گیرد چون دیگران دائماً در رقابت برنده می‌شوند. علل شامل قفل‌های ناعادلانه (unfair) زیر رقابت سنگین، سوءاستفاده از اولویت نخ، و یک writeLock پرمشغله که خواننده‌ها (reader) را گرسنه می‌کند (یا برعکس). راه‌حل‌ها:

  • جایی که دُم تأخیر (latency tail) مهم است از قفل‌های عادل (fair) استفاده کن: new ReentrantLock(true). قفل عادل مثل صفِ منظم نانوایی است — هرکس زودتر رسید زودتر می‌گیرد. انصاف، توان عملیاتی (throughput) را با انتظار کران‌دار معاوضه می‌کند، پس پیش از پیش‌فرض‌کردنش اندازه‌گیری کن.
  • برای بارهای خواندن‌محور از ReentrantReadWriteLock با سیاست انصاف/تنزل (downgrade)، یا StampedLock استفاده کن. توجه: StampedLock بازورودپذیر نیست و خواندن‌های خوش‌بینانه‌اش باید اعتبارسنجی شوند (کمی جلوتر می‌بینیمش).

شرایط رقابتی و بررسی-سپس-عمل (check-then-act)

رقابت (race condition) وقتی است که درستیِ برنامه به ترتیب درهم‌آمیزیِ نخ‌ها وابسته باشد — ترتیبی که تو کنترلش نمی‌کنی.

دو نفر و یخچال خالی

تو یخچال را باز می‌کنی، می‌بینی شیر تمام شده، و راه می‌افتی تا شیر بخری. هم‌خانه‌ات هم دقیقاً همان لحظه یخچال را دیده و او هم راه افتاده. نتیجه: دو نفر شیر می‌خرند. هر دو «بررسی» کردید (شیر نیست) و بعد «عمل» کردید (خرید)، اما بین بررسی و عمل، وضعیت زیر پایتان تغییر کرده بود. این دقیقاً الگوی بررسی-سپس-عمل (check-then-act) است.

رایج‌ترین شکل رقابت همین بررسی-سپس-عمل است: مقداری را مشاهده و بر اساسش عمل می‌کنی، اما مقدار در فاصلهٔ بین این دو تغییر می‌کند.

// باگ: بررسی-سپس-عمل کلاسیک — دو نخ می‌توانند هر دو از بررسی null عبور کنند
private Connection conn;
Connection get() {
    if (conn == null) {          // بررسی
        conn = open();           // عمل — دو اتصال نشت می‌کند، یا بدتر
    }
    return conn;
}

ConcurrentHashMap نسخهٔ ظریف‌تری از همین دام را دعوت می‌کند — ظریف چون خودِ نقشه نخ‌امن است و آدم گمان می‌کند خطر رفع شده:

// باگ: get سپس put اتمی نیست؛ دو نخ دوبار محاسبه می‌کنند، یکی برنده می‌شود
Value v = map.get(key);
if (v == null) {
    v = expensiveCompute(key);
    map.put(key, v);             // آخرین نوشتن برنده؛ کار هدررفته؛ v ناسازگار
}

نکته این‌جاست: هر عملیات ConcurrentHashMap به‌تنهایی اتمی است، اما وقتی یک get و یک put را کنار هم می‌گذاری، آن مجموعه دیگر اتمی نیست. راه‌حل، استفاده از یک عملیات اتمیِ واحد است — که دقیقاً همان دلیل وجود کالکشن‌های همزمان است:

// computeIfAbsent تابع نگاشت را به‌ازای هر کلید اتمی اجرا می‌کند
Value v = map.computeIfAbsent(key, this::expensiveCompute);
دام واقعی: `computeIfAbsent` بازگشتی

در جاوا ۸، فراخوانی بازگشتیِ computeIfAbsent روی همان نقشه برای کلیدی دیگر، داخل تابع نگاشت، می‌تواند جدول را خراب یا بن‌بست کند. جاوا ۹+ این تغییر بازورودی را تشخیص می‌دهد و استثنا پرتاب می‌کند. قاعدهٔ ساده: هرگز داخل تابع نگاشت، کارِ تغییردهندهٔ همان نقشه انجام نده.

حالا یک رقابت کوچک ولی همه‌جاحاضر — عملیات مرکب «خواندن-تغییر-نوشتن» روی یک شمارنده:

count++;                         // باگ: خواندن، افزودن، نوشتن — سه گام، نه اتمی

آن ++ بی‌گناه در واقع سه قدم است: مقدار را بخوان، یکی اضافه کن، بنویس. دو نخ می‌توانند همزمان مقدار قدیمی را بخوانند و هر دو همان مقدار به‌علاوهٔ یک را بنویسند — یک افزایش گم می‌شود. راه‌حل‌ها، به ترتیب صعودی مقیاس‌پذیری:

synchronized (lock) { count++; }             // درست، اما رقابتی
AtomicLong count = ...; count.incrementAndGet(); // CAS، بهتر زیر رقابت متوسط
LongAdder adder = ...; adder.increment();        // نواری، بهترین زیر رقابت بالا
چرا `LongAdder` زیر رقابت سنگین می‌بَرد

AtomicLong یک مکان حافظهٔ واحد دارد که همه سرش دعوا می‌کنند — مثل یک باجهٔ واحد که صف طولانی پشتش است. LongAdder کار را روی چند سلول (cell) پخش می‌کند — مثل بازکردن چند باجه — و فقط وقتی sum() را صدا می‌زنی همه را جمع می‌کند. پس وقتی نخ‌های زیادی می‌نویسند و تو به‌ندرت می‌خوانی، LongAdder برنده است. اما اگر مقدار را مدام می‌خوانی یا رقابت پایین است، AtomicLong ساده‌تر و کافی است.


تولیدکننده-مصرف‌کننده با BlockingQueue

پیشخوان آشپزخانه بین سالن و آشپزخانه

تصور کن یک پیشخوانِ محدود بین آشپزها (تولیدکننده) و پیشخدمت‌ها (مصرف‌کننده). آشپز غذا را روی پیشخوان می‌گذارد؛ پیشخدمت برمی‌دارد. اگر پیشخوان پر شود، آشپز مجبور است صبر کند (نه اینکه غذا را روی زمین تلنبار کند)؛ اگر خالی باشد، پیشخدمت صبر می‌کند. این «صبرِ خودکارِ دوطرفه» دقیقاً کاری است که BlockingQueue برایت می‌کند.

پیاده‌سازی دستی wait/notify برای تولیدکننده-مصرف‌کننده یک آیین گذار و منبع باگ‌های بی‌پایان است (سیگنال ازدست‌رفته، بیدارشدن گم‌شده، notify در برابر notifyAll). در محیط تولید تقریباً همیشه از BlockingQueue استفاده می‌کنی که بافر کران‌دار، انتظار شرطی، و فشار برگشتی (backpressure) را یکجا کپسوله می‌کند.

BlockingQueue<Task> queue = new ArrayBlockingQueue<>(1000); // کران‌دار → فشار برگشتی

// تولیدکننده
void produce(Task t) throws InterruptedException {
    queue.put(t);   // وقتی پر است مسدود می‌شود — این فشار برگشتیِ مطلوب است
}

// مصرف‌کننده با خاموش‌سازیِ قرص سمّی (poison pill)
static final Task POISON = new Task.Poison();
void consumeLoop() throws InterruptedException {
    while (true) {
        Task t = queue.take();        // وقتی خالی است مسدود می‌شود
        if (t == POISON) { queue.put(POISON); return; } // برای هم‌نوعان دوباره درج کن
        handle(t);
    }
}

«فشار برگشتی (backpressure)» اصطلاح مهمی است که همین‌جا بازش کنیم: یعنی وقتی مصرف‌کننده کند است، این کندی به عقب — به تولیدکننده — منتقل می‌شود و او را هم آرام می‌کند. بدون فشار برگشتی، تولیدکنندهٔ سریع، حافظه را پر می‌کند تا برنامه بترکد. انتخاب‌های کلیدی:

  • کران‌دار (ArrayBlockingQueue، LinkedBlockingQueue کران‌دار) فشار برگشتی می‌دهد — تولیدکننده‌ها کند می‌شوند به‌جای آنکه heap منفجر شود. کران‌دار را ترجیح بده. صف بی‌کران یک جهش بار را به OutOfMemoryError تبدیل می‌کند.
  • SynchronousQueue ظرفیت صفر دارد: هر put مستقیماً به یک take تحویل می‌دهد — مثل دست‌به‌دست کردن یک بشقاب داغ، نه گذاشتنش روی پیشخوان. این موتور پشت Executors.newCachedThreadPool است و یک ملاقات (rendezvous) واقعی را اجبار می‌کند.
  • قرص سمّی (poison pill) اصطلاح خاموش‌سازیِ تمیز است: یک نگهبان (sentinel) در صف بگذار تا مصرف‌کننده‌ها پس از تخلیهٔ کارهای واقعی خارج شوند، نه آنکه وسط کار قطع (interrupt) شوند. دوباره درجش کن تا چند مصرف‌کننده همه آن را ببینند.

نسخهٔ دستی (برای مصاحبه بدانش)

هرچند در عمل از BlockingQueue استفاده می‌کنی، مصاحبه‌گرها عاشق‌اند ببینند می‌توانی بافر کران‌دار را با دست بسازی. این نسخهٔ درست است:

// بافر کران‌دار درست با یک قفل و دو شرط (condition)
class BoundedBuffer<E> {
    private final Object[] buf;
    private int count, head, tail;
    private final ReentrantLock lock = new ReentrantLock();
    private final Condition notFull  = lock.newCondition();
    private final Condition notEmpty = lock.newCondition();

    BoundedBuffer(int cap) { buf = new Object[cap]; }

    void put(E e) throws InterruptedException {
        lock.lock();
        try {
            while (count == buf.length) notFull.await(); // while، نه if
            buf[tail] = e; tail = (tail + 1) % buf.length; count++;
            notEmpty.signal();
        } finally { lock.unlock(); }
    }
    @SuppressWarnings("unchecked")
    E take() throws InterruptedException {
        lock.lock();
        try {
            while (count == 0) notEmpty.await();
            E e = (E) buf[head]; buf[head] = null;       // برای GC null کن
            head = (head + 1) % buf.length; count--;
            notFull.signal();
            return e;
        } finally { lock.unlock(); }
    }
}

دو قاعدهٔ سطح‌ارشد در این کد تعبیه شده. اول: همیشه در یک while منتظر بمان، هرگز در یک if. چرا؟ چون بین لحظه‌ای که سیگنال می‌گیری و لحظه‌ای که قفل را دوباره می‌گیری، ممکن است نخ دیگری آن جای خالی را قاپیده باشد؛ و طبق مشخصات جاوا، «بیدارشدن کاذب (spurious wakeup)» هم مجاز است — یعنی گاهی بدون هیچ سیگنالی بیدار می‌شوی. تنها راه امن، بازبررسیِ شرط در یک حلقه است. دوم: از دو شرط جداگانه (notFull و notEmpty) استفاده کن تا یک signal روی «پر نیست» هرگز بیهوده یک مصرف‌کنندهٔ منتظرِ «خالی نیست» را بیدار نکند.

`if` به‌جای `while` = خرابیِ غیرقابل‌بازتولید

اگر به‌جای while از if استفاده کنی، نخی که کاذب یا دیرهنگام بیدار شده، بدون بررسی مجدد پیش می‌رود و روی وضعیت غلط عمل می‌کند — مثلاً از بافری که تازه دوباره پر شده برمی‌دارد. این نوع باگ فقط گاه‌به‌گاه و زیر بار خاص رخ می‌دهد و بازتولیدش تقریباً ناممکن است. با یک مانیتور واحد مجبور می‌شدی notifyAll بزنی که به اندازهٔ O(تعداد منتظران) هدررفته است.


اولیه‌های هماهنگی: latch، barrier، semaphore، phaser

تا این‌جا با قفل کار کردیم. اما گاهی نیاز نداری «دسترسی انحصاری» بدهی، بلکه می‌خواهی نخ‌ها را با هم هماهنگ کنی — مثلاً «همه با هم شروع کنید» یا «همه منتظر بمانید تا آخری برسد». برای این کار جاوا چند ابزار آماده دارد.

ابزارهای هماهنگی مثل ابزارهای یک مسابقهٔ دو

CountDownLatch مثل تپانچهٔ شروع است — یک‌بار شلیک می‌شود و همه می‌دوند. CyclicBarrier مثل خطی است که همهٔ دوندگان در پایان هر دور کنارش جمع می‌شوند و با هم دور بعد را شروع می‌کنند، بارها و بارها. Semaphore مثل تعداد محدود لاین استخر است — فقط N شناگر همزمان. Phaser مثل یک مربی منعطف که می‌تواند وسط تمرین دونده اضافه یا کم کند.

اولیه قابل‌استفادهٔ مجدد؟ کاربرد
CountDownLatch خیر (یک‌بارمصرف) انتظار برای تکمیل N رویداد پیش از ادامه
CyclicBarrier بله N نخ به‌طور مکرر در یک مانع ملاقات می‌کنند (محاسبهٔ فازی)
Semaphore بله محدودسازی دسترسی همزمان به N مجوز (استخر، سقف نرخ)
Phaser بله تعداد اعضای پویا، چندفازی؛ مانع منعطف
Exchanger بله دو نخ اشیاء را در یک ملاقات مبادله می‌کنند

بیا CountDownLatch را در عمل ببینیم — الگوی کلاسیک «همه را با هم شروع کن، بعد منتظر پایان همه بمان»:

// CountDownLatch: N کارگر را با هم شروع کن، منتظر پایان همه بمان
CountDownLatch ready = new CountDownLatch(1);   // دروازهٔ رهاسازی
CountDownLatch done  = new CountDownLatch(N);
for (int i = 0; i < N; i++) new Thread(() -> {
    ready.await();          // همه تا بازشدن دروازه مسدودند
    work();
    done.countDown();       // اعلام تکمیل
}).start();
ready.countDown();          // شلیک تپانچهٔ شروع
done.await();               // main منتظر همه می‌ماند

اینجا ready یک دروازه است که با یک شمارش تا صفر باز می‌شود؛ همهٔ کارگرها پشتش صبر می‌کنند تا main تپانچه را بزند. done از N شروع می‌شود و هر کارگر با پایان کارش یکی کم می‌کند؛ وقتی به صفر رسید، main رها می‌شود.

CountDownLatch تا صفر می‌شمارد و همان‌جا می‌ماند — قابل بازنشانی (reset) نیست، یک‌بارمصرف است. وقتی به ملاقات تکرارپذیر نیاز داری، از CyclicBarrier استفاده کن که می‌تواند هنگام رسیدن آخرین نخ یک اقدام مانع (barrier action) اجرا کند و سپس خودش را بازنشانی کند:

CyclicBarrier barrier = new CyclicBarrier(N, () -> mergePhaseResults());
// هر کارگر در پایان هر فاز barrier.await() را صدا می‌زند
`CyclicBarrier` همه-یا-هیچ است

اگر یکی از نخ‌های منتظر قطع (interrupt) شود یا مهلتش تمام شود، مانع شکسته (broken) می‌شود و هر منتظر دیگری یک BrokenBarrierException می‌گیرد — نه فقط آن یک نخ. مثل تیم کوهنوردی که با طناب به هم بسته‌اند: اگر یکی بیفتد، همه را می‌کشد. کد مقاوم باید این استثنا را بگیرد و فاز را تمیز بازنشانی یا رد کند.

حالا Semaphore که همزمانی را کران‌دار می‌کند — همان استخر اتصال (connection pool) یا محدودکنندهٔ نرخِ متعارف:

Semaphore permits = new Semaphore(10, /*fair*/ true);
void call() throws InterruptedException {
    permits.acquire();
    try { doRemoteCall(); } finally { permits.release(); } // همیشه در finally آزاد کن
}
باگ شمارهٔ یکِ سمافور: آزادکردن بیش از حد

release() در برابر acquire()ِ پیشین اعتبارسنجی نمی‌شود. اگر یک release() اضافه در یک مسیر کد داشته باشی، بی‌سروصدا شمار مجوزها را بالا می‌برد — سمافور از ۱۰ مجوز به ۱۱ و بیشتر می‌رسد و کل محدودیت همزمانی نابود می‌شود، بی‌آنکه هیچ خطایی ببینی. قاعده: acquire را بیرون از try بگیر، و در finally دقیقاً یک‌بار release کن.

و در آخر Phaser: هم latch و هم barrier را تعمیم می‌دهد. اعضا می‌توانند به‌صورت پویا ثبت/لغو ثبت (register/deregister) شوند و از چند فاز بدون بازسازی پشتیبانی می‌کند — ایده‌آل برای خط‌لوله‌های مرحله‌ای به‌سبک fork/join که تعداد شرکت‌کنندگان در طول اجرا تغییر می‌کند.


سه مسئلهٔ کلاسیک

هر کتاب همزمانی سه معمای مشهور دارد که هرکدام یکی از شرط‌های کافمن را روشن می‌کنند. اولی را از قبل حل کردیم.

بافر کران‌دار — بالا حل شد (تولیدکننده-مصرف‌کننده).

خواننده‌ها-نویسنده‌ها (readers–writers)

یک تخته‌اعلانات

تصور کن یک تخته‌اعلانات: هزار نفر می‌توانند همزمان بخوانند بدون هیچ مشکلی، اما وقتی یک نفر می‌خواهد چیزی بنویسد یا پاک کند، همه باید کنار بروند تا او تنها بماند. خواندن اشتراکی است، نوشتن انحصاری.

خواننده‌های زیادی می‌توانند اشتراک داشته باشند؛ یک نویسنده به انحصار نیاز دارد. ReentrantReadWriteLock این را مدیریت می‌کند، اما استفادهٔ ساده‌لوحانه زیر ترافیک خواندن مداوم، نویسنده‌ها را گرسنه می‌کند (اگر همیشه یک خواننده در حال خواندن باشد، نویسنده هرگز نوبت نمی‌گیرد).

ReentrantReadWriteLock rw = new ReentrantReadWriteLock(true); // عادل → بدون گرسنگی نویسنده
Lock r = rw.readLock(), w = rw.writeLock();

Object read()  { r.lock(); try { return data; } finally { r.unlock(); } }
void  write(Object x) { w.lock(); try { data = x; } finally { w.unlock(); } }
می‌توانی تنزل بدهی، اما نمی‌توانی ارتقا بدهی

می‌توانی تنزل (downgrade) دهی: قفل نوشتن را نگه دار، قفل خواندن را بگیر، بعد نوشتن را آزاد کن — این امن است. اما نمی‌توانی ارتقا (upgrade) دهی: قفل خواندن را نگه داری و بخواهی نوشتن بگیری — چون نویسنده باید منتظر رفتن همهٔ خواننده‌ها بماند، از جمله خودت که هنوز قفل خواندن را رها نکرده‌ای. نتیجه خودبن‌بستی است.

برای دادهٔ خواندن‌غالب، StampedLock یک ترفند فوق‌العاده دارد: خواندن خوش‌بینانه (optimistic read) که در مسیر خوشحال (happy path) اصلاً قفلی نمی‌گیرد.

StampedLock sl = new StampedLock();
double distanceFromOrigin() {
    long stamp = sl.tryOptimisticRead();      // قفلی گرفته نمی‌شود
    double cx = x, cy = y;                     // فیلدها را بخوان
    if (!sl.validate(stamp)) {                 // نویسنده‌ای مداخله کرد؟
        stamp = sl.readLock();                 // به قفل خواندن واقعی برگرد
        try { cx = x; cy = y; } finally { sl.unlockRead(stamp); }
    }
    return Math.sqrt(cx * cx + cy * cy);
}

منطقش این است: یک «مُهر (stamp)» بگیر، بدون قفل بخوان، بعد بپرس «آیا در این فاصله نویسنده‌ای آمد؟». اگر نه، خواندنت معتبر بود و رایگان تمام شد. اگر بله، به یک قفل خواندن واقعی برگرد. توجه: StampedLock بازورودپذیر نیست و از Condition پشتیبانی نمی‌کند — این محدودیت‌ها را رعایت کن.

فیلسوفان شام‌خوار (dining philosophers)

میز شام فیلسوف‌ها

پنج فیلسوف دور یک میز گرد نشسته‌اند و بین هر دو نفر یک چنگال است — جمعاً پنج چنگال. هر فیلسوف برای غذاخوردن به هر دو چنگال کنارش نیاز دارد. اگر همه همزمان چنگال چپشان را بردارند، هرکس یک چنگال دارد و منتظر چنگال راست است که دست همسایه است — و همه گرسنه می‌مانند. این یک نمایش زندهٔ انتظار دوری است.

«چنگال چپ را بگیر، سپس راست» ساده‌لوحانه، وقتی همه همزمان چپ را بگیرند بن‌بست می‌کند. دو راه‌حل تمیز داریم که هرکدام یک شرط کافمن متفاوت را می‌شکنند:

// راه‌حل A: تقارن را بشکن — یک فیلسوف اول راست را بردارد (ترتیب منابع)
void dine(int id, Lock left, Lock right) {
    Lock first = (id == LAST) ? right : left;   // یک فیلسوف معکوس می‌کند
    Lock second = (id == LAST) ? left : right;
    first.lock();
    try { second.lock();
        try { eat(); } finally { second.unlock(); }
    } finally { first.unlock(); }
}
// راه‌حل B: با یک سمافور، همزمانی را به N-۱ فیلسوف نشسته محدود کن
Semaphore seats = new Semaphore(PHILOSOPHERS - 1);
void dine(...) throws InterruptedException {
    seats.acquire();               // حداکثر ۴ از ۵ می‌توانند برای چنگال‌ها رقابت کنند
    try { left.lock(); right.lock();
        try { eat(); } finally { right.unlock(); left.unlock(); }
    } finally { seats.release(); }
}
دو راه‌حل، دو شرط شکسته‌شده

راه‌حل A با معکوس‌کردنِ ترتیب یک فیلسوف، انتظار دوری را می‌شکند (ترتیب سراسری / عدم تقارن). راه‌حل B با اجازه‌دادن حداکثر به N-۱ نفر، نگه‌داشتن و انتظار را می‌شکند — چون با ۴ نفر برای ۵ چنگال، دست‌کم یک نفر همیشه می‌تواند هر دو چنگالش را بگیرد و بن‌بست از نظر ریاضی ناممکن می‌شود. این نشان می‌دهد چطور شکستن هر یک از چهار شرط کافمن کافی است.


تغییرناپذیری (immutability) و محصورسازی (confinement) به‌عنوان راهبرد

تا این‌جا کلی ابزار برای مدیریت اشتراک یاد گرفتیم. اما بهترین راهبرد این است که اصلاً اشتراکِ تغییرپذیر نداشته باشی.

ارزان‌ترین همزمانی، نبودِ وضعیت مشترک تغییرپذیر است

اگر داده‌ای مشترک اما تغییرناپذیر باشد، یا اصلاً مشترک نباشد، هیچ قفلی لازم نیست و هیچ رقابتی ممکن نیست. دو راهبرد ساختاری — تغییرناپذیری و محصورسازی — کلاس‌های کاملی از باگ را از ریشه حذف می‌کنند.

تغییرناپذیری (immutability).

یک سنگِ حکاکی‌شده

یک شیء تغییرناپذیر مثل یک لوح سنگی حکاکی‌شده است: وقتی ساخته شد، دیگر هیچ‌کس نمی‌تواند تغییرش دهد. هزار نفر می‌توانند همزمان بخوانندش بی‌هیچ خطری، چون هیچ‌کس نمی‌نویسد. اگر چیز جدیدی خواستی، یک لوح تازه می‌سازی.

شیئی که همهٔ فیلدهایش final هستند و هرگز حین ساخت فرار نمی‌کنند، از طریق تضمین فیلد-final در JMM به‌طور ایمن منتشر (safely published) می‌شود و می‌تواند بدون همگام‌سازی آزادانه به‌اشتراک گذاشته شود. recordهای جاوا این را طبیعی می‌کنند:

record Money(long cents, String currency) {          // عمیقاً تغییرناپذیر
    Money add(Money o) {                              // یک نمونهٔ جدید برمی‌گرداند
        if (!currency.equals(o.currency)) throw new IllegalArgumentException();
        return new Money(cents + o.cents, currency);
    }
}
`final` از ارجاع محافظت می‌کند، نه از شیءِ پشتش

یک record که یک List نگه می‌دارد، تغییرناپذیر نیست مگر آنکه دفاعی (defensive copy) آن را به یک لیست تغییرناپذیر کپی کنی — وگرنه کسی می‌تواند محتوای لیست را عوض کند هرچند خودِ ارجاع final است. و یک فیلد final فقط زمانی ایمن منتشر می‌شود که this حین اجرای سازنده فرار نکرده باشد (مثلاً خودت را در یک شنونده ثبت نکرده باشی پیش از پایان ساخت).

محصورسازی (confinement). داده را روی یک نخ نگه دار تا اصلاً به همگام‌سازی نیاز نباشد. سه شکل دارد:

  • محصورسازی نخی (thread confinement) با ThreadLocal — هر نخ نسخهٔ خودش را دارد. اما مراقب نشت باش: روی یک استخر نخ (thread pool)، ThreadLocalی که remove() نکنی تا زمان زنده‌بودن نخ کارگر می‌ماند و می‌تواند اشیاء بزرگ یا کلاس‌لودرها را پین کند — یک نشت حافظهٔ کلاسیک در وب‌اپ‌ها. همیشه در یک finally، remove() کن.
  • محصورسازی پشته‌ای (stack confinement) — متغیرها و اشیاء محلی که هرگز از یک متد فرار نمی‌کنند خودبه‌خود نخ‌امن‌اند، چون هر نخ پشتهٔ خودش را دارد. این را به‌طور پیش‌فرض ترجیح بده؛ رایگان است.
  • محصورسازی نمونه‌ای (instance confinement) — وضعیت تغییرپذیر را پشت قفل خودِ یک شیء محافظت کن و هرگز نگذار ارجاعی فرار کند. این همان الگوی مانیتور جاوا است که عمداً و آگاهانه انجام شده.
نخ‌های مجازی (جاوا ۲۱) قواعد را عوض نمی‌کنند

نخ‌های مجازی (virtual threads، جاوا ۲۱، Thread.ofVirtual()) مسدودشدن را ارزان می‌کنند تا بتوانی کد مسدودشوندهٔ سرراست و محصورشدهٔ به‌ازای هر درخواست را در مقیاس عظیم بنویسی. اما یک شیء مشترک تغییرپذیر از یک نخ مجازی دقیقاً به همان اندازهٔ یک نخ پلتفرمی ناامن است و قواعد JMM دست‌نخورده‌اند. توجه: اگر یک نخ مجازی داخل یک بلوک synchronized مسدود شود، پین‌شدن (pinning) رخ می‌دهد؛ روی جاوا ۲۱ به این دلیل ReentrantLock را ترجیح بده (این محدودیت تا حد زیادی در جاوا ۲۴+ حل شده است). و هرگز نخ‌های مجازی را استخر (pool) نکن.


نکات آزمون همزمانی

نمی‌توانی درستی را با آزمون *وارد* یک طراحی همزمان کنی

باگ‌های همزمانی نامعین (non-deterministic) هستند؛ یک آزمون واحد ممکن است هزار بار سبز شود و بار هزار و یکم زیر بار واقعی بترکد. آزمون‌های عادی اعتماد کاذب می‌دهند.

تکنیک‌هایی که واقعاً این باگ‌ها را پیدا می‌کنند:

  • jcstress — بستر آزمون OpenJDK که مشخصاً برای آشکارکردن باگ‌های JMM/بازچینش ساخته شده، با اجرای میلیاردها درهم‌آمیزی و طبقه‌بندی نتایج. برای هر کد بدون‌قفلِ سطح‌پایینی که می‌نویسی از آن استفاده کن.
  • فشار با رقابت (contention): نخ‌های فراوان (بیش از تعداد هسته‌ها) را در یک حلقهٔ فشرده برای چند ثانیه اجرا کن، از یک CyclicBarrier استفاده کن تا همه دقیقاً در یک لحظه شروع کنند (بیشینه‌کردن هم‌پوشانی)، و پس از آن یک ناوردا (invariant) را ادعا کن.
  • -Xint / -XX:-TieredCompilation و اجرا روی سخت‌افزار ARM/حافظه‌ضعیف بازچینشی را آشکار می‌کند که مدل حافظهٔ قوی x86 پنهانش می‌کند.
  • نگهبان بن‌بست در آزمون‌ها: یک نظرسنجی پس‌زمینهٔ ThreadMXBean.findDeadlockedThreads() که لحظهٔ ظاهرشدن یک چرخه، آزمون را رد می‌کند، به‌جای معلق‌کردن CI.
  • Thread.sleep در آزمون‌ها بوی بد می‌دهد — آزمون‌ها را کند و همچنان بی‌ثبات (flaky) می‌کند. از latch/barrier برای بیان ترتیب واقعی‌ای که می‌خواهی تحمیل کنی استفاده کن.
  • زمان‌بند را فاز کن (fuzz): ابزارهایی مانند تزریق Thread.yield()، یا وارسی مدل با JPF (Java PathFinder) برای درهم‌آمیزی‌های جامعِ دامنه‌کوچک.

حقیقت صادقانه: درستی را با happens-before استدلال می‌کنی، سطح مشترک را کوچک نگه می‌داری، و از آزمون‌ها فقط برای گرفتن پس‌رفت‌ها (regression) استفاده می‌کنی.


سؤالات مصاحبه

حالا وقت آن است که همه‌چیز را در قالب سؤال‌های واقعی مصاحبه جمع کنیم. هر سؤال را با پاسخ کامل بخوان و سعی کن پیش از دیدن پاسخ، خودت جواب بدهی.

۱) چهار شرط کافمن را بیان کن و برای هرکدام یک راه پیشگیری بده.

انحصار متقابل (تغییرناپذیر/بدون‌قفل)، نگه‌داشتن-و-انتظار (همهٔ قفل‌ها را یکجا بگیر)، عدم پیش‌دستی (tryLock + مهلت)، انتظار دوری (ترتیب سراسری قفل). شکستن هر یک، بن‌بست را دفع می‌کند؛ حذف انتظار دوری از طریق ترتیب‌دهی قفل در بیشتر سیستم‌ها عملی‌ترین است.

۲) تفاوت بن‌بست، زنده‌قفلی و گرسنگی.

بن‌بست: نخ‌ها برای همیشه در یک چرخهٔ انتظار مسدودند. زنده‌قفلی: نخ‌ها فعالانه در حال اجرا و واکنش‌اند اما هیچ پیشرفتی ندارند (تلاش‌مجدد متقارن). گرسنگی: یک نخ پیشرفت نمی‌کند چون دیگران مدام منبع را می‌برند. بن‌بست/زنده‌قفلی معمولاً مسائل تقارن/ترتیب‌اند؛ گرسنگی یک مسئلهٔ انصاف است.

۳) (دام) چرا انتظار روی شرط باید `while` باشد، نه `if`؟

بیدارشدن‌های کاذب (spurious wakeup) طبق مشخصات مجازند، و حتی بدون آن‌ها، نخ دیگری ممکن است بین بیدارشدن تو و بازگرفتن قفل، شرط را مصرف کند. بازبررسیِ گزارهٔ (predicate) در یک حلقه تنها الگوی درست است. یک if خرابیِ گاه‌به‌گاه تولید می‌کند که بازتولیدش تقریباً ناممکن است.

۴) `notify` در برابر `notifyAll` — چه زمانی `notify` امن است؟

notify یک منتظرِ دلخواه را بیدار می‌کند. تنها زمانی امن است که همهٔ منتظران قابل‌تعویض باشند (روی یک شرط منتظرند و پیشرفت هرکدام کافی است) و تو دقیقاً یک واحد موجود را سیگنال دهی. اگر منتظران روی گزاره‌های متفاوتِ همان مانیتور منتظر باشند، notify می‌تواند نادرست را بیدار کند و باعث معلق‌شدنِ «بیدارشدن-گم‌شده» شود؛ از notifyAll یا بهتر، Conditionهای مجزا استفاده کن.

۵) (دام) این چه چاپ می‌کند؟
List<Integer> list = new ArrayList<>();
IntStream.range(0, 4).parallel().forEach(list::add);
System.out.println(list.size());

نامعین — هرچیزی از 1 تا 4، یا یک ArrayIndexOutOfBoundsException/NullPointerException. ArrayList نخ‌امن نیست؛ add همزمان روی size و آرایهٔ پشتیبان رقابت می‌کند. اصلاح: Collections.synchronizedList، یک کالکشن همزمان، یا .collect(Collectors.toList()) روی استریم.

۶) (دام — باگ را پیدا کن)
if (!map.containsKey(k)) map.put(k, compute(k)); // map یک ConcurrentHashMap است

رقابت بررسی-سپس-عمل: دو نخ هر دو کلید را غایب می‌بینند و هر دو محاسبه/درج می‌کنند. عملیات جداگانهٔ ConcurrentHashMap اتمی‌اند، اما عملیات مرکب اتمی نیست. اصلاح: map.computeIfAbsent(k, this::compute).

۷) چه زمانی `LongAdder` را بر `AtomicLong` ترجیح می‌دهی؟

رقابت نوشتنِ بالا با خواندن‌های کم‌تکرار. مکان تک-CAS در AtomicLong یک نقطهٔ داغ می‌شود؛ LongAdder روی سلول‌ها نوار می‌کشد و هنگام خواندن جمع می‌زند، و خواندنِ همیشه-دقیق و حافظه را با توان عملیاتی نوشتنِ بسیار بالاتر معاوضه می‌کند. برای رقابت پایین یا وقتی مقدار را مدام می‌خوانی، AtomicLong ساده‌تر و خوب است.

۸) چرا BlockingQueue کران‌دار را بر بی‌کران ترجیح می‌دهی؟

فشار برگشتی (backpressure). صف کران‌دار وقتی پر است تولیدکننده‌ها را مسدود می‌کند و کندی را به بالادست منتشر می‌کند. صف بی‌کران یک جهش بار را در رشد نامحدودِ heap جذب می‌کند و سرانجام OutOfMemoryError می‌دهد، و یک مسئلهٔ تأخیر را به یک قطعی (outage) تبدیل می‌کند. ظرفیت یک پارامتر طراحی است، نه مزاحمت.

۹) (سخت) تضمین فیلد-`final` در JMM را توضیح بده و اینکه چگونه یک رقابت داده‌ای باز هم می‌تواند بشکندش.

اگر فیلدهای final یک شیء در سازنده مقداردهی شوند و this حین ساخت فرار نکند، هر نخی که ارجاعی به شیء ببیند تضمین می‌شود فیلدهای finalِ درست-مقداردهی‌شده را بدون همگام‌سازی ببیند. وقتی می‌شکند که this از سازنده فرار کند (مثلاً ثبت یک شنونده پیش از پایان ساخت) — آنگاه نخ دیگری می‌تواند وضعیت نیمه‌ساخته را مشاهده کند.

۱۰) آیا دو نخ می‌توانند با یک قفل واحد بن‌بست کنند؟

روی یک قفل بازورودپذیر که همان نخ دوباره بگیرد، نه. اما بله بین دو نخ اگر یکی قفل A را نگه دارد و متدی را صدا بزند که به قفل B نیاز دارد در حالی که دیگری B را نگه دارد و A را می‌خواهد — «شیء واحد» باز می‌تواند بخشی از یک چرخهٔ دوقفلی باشد. همچنین یک قفل غیربازورودپذیر (مانند StampedLock) می‌تواند خودبن‌بستی کند اگر همان نخ دوباره قفلش کند.

۱۱) (دام) یک `CyclicBarrier(3)` دو نخ منتظر دارد و نخ سوم روی `await` مهلتش تمام می‌شود. بر سر دو منتظر چه می‌آید؟

یک BrokenBarrierException می‌گیرند. مانع همه-یا-هیچ است: یک مهلت، قطع (interrupt)، یا اقدام ناموفق آن را برای همهٔ منتظران فعلی می‌شکند. کد مقاوم این را می‌گیرد و فاز را تمیز بازنشانی یا رد می‌کند.

۱۲) دام درستی سمافور؟

release() در برابر acquire()ِ پیشین اعتبارسنجی نمی‌شود. یک release اضافه — اغلب در مسیر استثنایی که در یک finally بدون acquire()ِ متناظر هم اجرا شده — بی‌سروصدا شمار مجوز را بالا می‌برد و حد همزمانی را نابود می‌کند. بیرون از try بگیر، در finally دقیقاً یک‌بار آزاد کن.

۱۳) (سخت) چرا `ThreadLocal` روی یک استخر نخ نشت می‌کند و چگونه پیشگیری می‌کنی؟

نخ‌های استخر عمرِ بلند دارند، پس مقداری که set می‌کنی در وظایف نامرتبط باقی می‌ماند و هرچه به آن ارجاع دارد را پین می‌کند (بافرهای بزرگ، کلاس‌لودرها در سرورهای اپلیکیشن) تا نخ بمیرد. با پیچیدن استفاده در try/finally { threadLocal.remove(); } در مرز وظیفه پیشگیری کن. InheritableThreadLocal این را در نخ‌های زاده‌شده تشدید می‌کند.

۱۴) نخ‌های مجازی (جاوا ۲۱) طراحی همزمانی را چگونه تغییر می‌دهند و چه چیزی ثابت می‌ماند؟

مسدودشدن را ارزان می‌کنند، پس می‌توانی از کد ساده و همگام و نخ-به‌ازای-هر-درخواست و محصور در مقیاس استفاده کنی به‌جای زنجیره‌های فراخوانِ واکنشی — استخرهای نخ کمتری برای تنظیم. چه چیزی ثابت می‌ماند: وضعیت مشترک تغییرپذیر دقیقاً همان‌قدر خطرناک است، و قواعد JMM بدون تغییرند. مراقب پین‌شدن هنگام مسدودشدن داخل synchronized باش (روی ۲۱ ReentrantLock را ترجیح بده) و هرگز نخ‌های مجازی را استخر نکن.

۱۵) (سخت) سناریویی بده که افزودن همگام‌سازی *یک بن‌بست بسازد*، و چگونه تشخیصش می‌دهی.

دو متدِ پیش‌تر مستقل را در synchronized می‌پیچی تا یک رقابت را رفع کنی؛ حالا یک گراف فراخوانی روی یک مسیر قفل A→B و روی مسیر دیگر B→A می‌گیرد و یک چرخه می‌سازد. با یک thread dump (jstack بن‌بست و چرخه را چاپ می‌کند) یا ThreadMXBean.findDeadlockedThreads() در یک نگهبان تشخیص بده. با تحمیل ترتیب سراسری قفل یا کوچک‌کردن ناحیهٔ بحرانی تا قفل‌گذاری تودرتو حذف شود، اصلاح کن.


نکاتِ سنیور و موارد پیشرفته

تا اینجا نقشهٔ باگ‌های همزمانی و الگوهای دفاعی را ساختیم. اما چیزی که یک سنیورِ واقعی را از یک برنامه‌نویسِ خوب جدا می‌کند، دانستنِ synchronized و BlockingQueue نیست — این‌ها را همه بلدند. تفاوت در جاهایی است که کد در پروداکشن زیر بار واقعی می‌شکند: استخر نخ‌ها که بی‌صدا نخ‌های اضافه‌اش را نادیده می‌گیرد، InterruptedException که کسی بلعیده و کل سیستم دیگر cancel نمی‌شود، CompletableFuture که روی یک استخرِ اشتباه بلاک شده، و باگ‌هایی که در سطحِ خطِ کش پردازنده زندگی می‌کنند. این بخش دقیقاً همین لایه است.

نقشهٔ این بخش

اول به قلبِ تپندهٔ هر سرویسِ جاوا می‌رویم: ThreadPoolExecutor و تلهٔ صف نامحدودش. بعد پروتکل interruption را یاد می‌گیریم (چرا «بلعیدنِ InterruptedException» یک جرم است). سپس CompletableFuture و تلهٔ commonPool و الگوی Memoizer برای شکستِ cache stampede. بعد double-checked locking و اصطلاحِ holder. سپس به سختِ‌افزار می‌رسیم: false sharing. بعد مسئلهٔ ABA در CAS. و آخر، نقشهٔ همزمانیِ جاوای مدرن (۲۰۲۵–۲۰۲۶): مرگِ biased locking، حلِ pinning در JDK 24، و Scoped Values. بعد چند سؤالِ سختِ سنیور.

استخر نخ‌ها؛ جایی که بیشترِ حوادثِ پروداکشن زاده می‌شوند

بیشترِ کدِ همزمانی‌ای که در عمل می‌نویسی، مستقیم با نخ کار نمی‌کند — یک ExecutorService می‌سازی و کارها را به آن می‌سپاری. پشتِ Executors.newFixedThreadPool(...) یک ThreadPoolExecutor نشسته و همین کلاس یک منطقِ پذیرشِ کار دارد که اگر ندانی، یک روز غافلگیرت می‌کند.

رستورانی با آشپز، صندلی انتظار و پیشخدمتِ اضافه

corePoolSize تعدادِ آشپزهای همیشه‌سرِ‌کار است. queue صندلی‌های انتظارِ سفارش‌هاست. maximumPoolSize سقفِ آشپزهایی که در اوجِ شلوغی می‌توانی صدا بزنی. قانونِ طلایی این است: تازه وقتی صف پُر شد، آشپزِ اضافه استخدام می‌شود — نه زودتر.

ترتیبِ پذیرشِ یک تسکِ جدید دقیقاً این است و ترتیبش همان‌جایی است که همه اشتباه می‌کنند:

۱) اگر نخ‌های فعال < corePoolSize  → یک نخِ هسته بساز و اجرا کن
۲) وگرنه، تسک را در صف بگذار (queue.offer)
۳) اگر صف پُر بود  → تا maximumPoolSize نخِ اضافه بساز
۴) اگر آن هم پُر بود → RejectedExecutionHandler را صدا بزن

نمودارِ زیر همین جریانِ تصمیم را نشان می‌دهد (Task admission flow):

flowchart TD
  A[New task submitted] --> B{active < corePoolSize?}
  B -- yes --> C[Start core thread]
  B -- no --> D{queue.offer succeeds?}
  D -- yes --> E[Task waits in queue]
  D -- no --> F{active < maximumPoolSize?}
  F -- yes --> G[Start extra thread]
  F -- no --> H[RejectedExecutionHandler]

حالا تلهٔ کشنده: قدم ۲ می‌گوید «تا وقتی صف جا دارد، در صف بگذار». اگر صفت نامحدود باشد (مثل LinkedBlockingQueue بدون ظرفیت، که دقیقاً همان چیزی است که newFixedThreadPool می‌سازد)، offer هیچ‌وقت شکست نمی‌خورد. یعنی قدم ۳ هرگز اجرا نمی‌شود و maximumPoolSize تو صرفاً یک عددِ تزئینی است.

`maximumPoolSize` با صفِ نامحدود یک دروغِ کامل است

اگر یک ThreadPoolExecutor با corePoolSize=10, maximumPoolSize=100 و یک LinkedBlockingQueue بی‌ظرفیت بسازی، سیستمت هیچ‌وقت از ۱۰ نخ فراتر نمی‌رود، هرچقدر هم بار بیاید. کارها بی‌صدا در صف تلنبار می‌شوند، تأخیر منفجر می‌شود، و در نهایت heap پُر و OutOfMemoryError. این دقیقاً همان دلیلی است که تیم‌ها به‌جای Executors.* مستقیم new ThreadPoolExecutor(...) با یک صفِ محدود و یک saturation policy می‌سازند.

سیاست‌های اشباع (وقتی هم صف و هم نخ‌ها پُرند) را باید عمداً انتخاب کنی:

  • AbortPolicy (پیش‌فرض): RejectedExecutionException پرتاب می‌کند — کالر خبردار می‌شود.
  • CallerRunsPolicy: تسک را در همان نخِ فراخوان اجرا می‌کند. این یک ترمزِ backpressureِ خودکارِ نابغه‌آسا است — نخِ وب که باید تسک بدهد، خودش مجبور می‌شود کار کند و در آن مدت تسکِ جدیدی نمی‌پذیرد.
  • DiscardPolicy / DiscardOldestPolicy: بی‌صدا دور می‌ریزد — تقریباً همیشه اشتباه، چون بی‌سروصدا داده گم می‌کنی.
فرمولِ سایزینگِ استخر (Brian Goetz)

برای بارِ محاسباتی‌محور: تعدادِ نخ ≈ تعدادِ هسته + ۱. برای بارِ I/O‑محور، فرمولِ کلاسیک این است: N = N_cpu × U × (1 + W/C) که U بهره‌وریِ هدف (۰ تا ۱)، W زمانِ انتظار (I/O) و C زمانِ محاسبهٔ هر تسک است. نکتهٔ سنیور: هرچه نسبتِ W/C بزرگ‌تر (کارِ بیشترِ I/O)، به نخِ بیشتری نیاز داری. اما با ورودِ virtual threadها این حساب‌وکتاب برای کارِ I/O‑محور تقریباً منسوخ شد — یک Executors.newVirtualThreadPerTaskExecutor() می‌سازی و دیگر سایزِ استخر را تیون نمی‌کنی.

و دو نکتهٔ خاموش‌سازی که در ریویو مدام می‌بینم اشتباه است:

`shutdown()` منتظر نمی‌ماند و `Future.get` استثنا را می‌پیچد

executor.shutdown() فقط «دیگر تسکِ جدید نپذیر» را علامت می‌زند و بلافاصله برمی‌گردد؛ برای اینکه واقعاً صبر کنی باید بعدش awaitTermination(...) صدا بزنی. shutdownNow() هم نخ‌ها را interrupt می‌کند اما فقط اگر کدت به interrupt احترام بگذارد (بخشِ بعد). و وقتی تسک استثنا پرتاب کند، future.get() آن را در یک ExecutionException می‌پیچد — باید e.getCause() را باز کنی. بدتر: اگر با execute(...) (نه submit) تسک بدهی و استثنا رخ دهد، استثنا بی‌صدا به UncaughtExceptionHandler می‌رود و تو در لاگ چیزی نمی‌بینی مگر آن را ست کرده باشی.

interruption یک پروتکل است، نه یک کلیدِ کشتن

بزرگ‌ترین سوءتفاهمِ سنیورهای نوپا این است که فکر می‌کنند thread.interrupt() نخ را «می‌کشد». نه. interrupt فقط یک پرچمِ boolean را روی نخ روشن می‌کند؛ این یک درخواستِ مؤدبانهٔ لغو است که خودِ کد باید به آن پاسخ دهد. کلِ مکانیزمِ cancellation در جاوا روی همین قرارداد بنا شده.

جنایتِ شمارهٔ یک: بلعیدنِ `InterruptedException`

این را همه‌جا می‌بینی:

try {
    Thread.sleep(1000);
} catch (InterruptedException e) {
    // ... هیچ
}

وقتی متدی که پرتابش می‌کند InterruptedException گرفت، JVM پرچمِ interrupt را پاک می‌کند. اگر تو استثنا را بگیری و کاری نکنی، سیگنالِ لغو برای همیشه ناپدید شد — لایه‌های بالاتر دیگر نمی‌فهمند این نخ باید بمیرد، و یک shutdownNow() یا یک timeout بی‌اثر می‌شود. نتیجه: نخ‌هایی که هیچ‌وقت تمام نمی‌شوند و در thread dump تلنبار شده‌اند.

قانونِ درست دو حالت دارد. اگر می‌توانی InterruptedException را به بالا propagate کنی، بکن (بگذار متدت آن را throw کند). اگر نمی‌توانی (مثلاً در یک Runnable که امضایش اجازه نمی‌دهد)، باید پرچم را دوباره روشن کنی:

try {
    queue.take();
} catch (InterruptedException e) {
    Thread.currentThread().interrupt(); // پرچم را برگردان
    return; // و از حلقه/تسک بیرون بزن
}

و برای حلقه‌های محاسباتیِ طولانی که هیچ متدِ بلاک‌کننده‌ای ندارند (پس InterruptedException نمی‌گیرند)، باید خودت پرچم را چک کنی:

while (!Thread.currentThread().isInterrupted()) {
    doOneChunkOfWork();
}
interruption قابلیتِ همکاری (cooperative) است

هیچ‌کس نخِ تو را با زور متوقف نمی‌کند (Thread.stop() سال‌هاست deprecated و خطرناک است چون قفل‌ها را در وسطِ کار رها می‌کند). لغو فقط وقتی کار می‌کند که همهٔ کدِ مسیر — کدِ تو و کتابخانه‌هایت — به پرچمِ interrupt احترام بگذارند. یک کتابخانهٔ بد که استثنا را می‌بلعد، cancellation را برای کلِ سیستم می‌شکند.

CompletableFuture: ترکیبِ ناهمگام و تله‌هایش

فصلِ اصلی از Future و استخرها گفت اما از ابزارِ ترکیبِ ناهمگامِ مدرن — CompletableFuture (از Java 8) — نگفت. این همان چیزی است که به تو اجازه می‌دهد به‌جای بلاک‌کردن روی future.get()، یک زنجیرهٔ غیرمسدودکننده از مراحل بسازی.

اولین تمایزی که در مصاحبه می‌پرسند: thenApply در برابر thenCompose.

// thenApply: تابعِ همگام، مقدار → مقدار
CompletableFuture<Integer> len = fetchUser(id).thenApply(User::name).thenApply(String::length);

// thenCompose: تابعی که خودش CompletableFuture برمی‌گرداند → صاف می‌کند (flatMap)
CompletableFuture<Order> order = fetchUser(id).thenCompose(u -> fetchLatestOrder(u)); // نه فیوچرِ تودرتو

اگر با thenApply تابعی بدهی که خودش CompletableFuture برمی‌گرداند، به CompletableFuture<CompletableFuture<T>> می‌رسی — دقیقاً مثلِ map در برابر flatMap در استریم. thenCompose همان flatMap است.

تلهٔ خاموش: `commonPool` و متدهای `*Async`

هر متدِ بدونِ پسوندِ Async روی همان نخی اجرا می‌شود که مرحلهٔ قبل را کامل کرده. متدهای *Async بدونِ آرگومانِ Executor روی ForkJoinPool.commonPool() اجرا می‌شوند. مشکل: commonPool به‌طور پیش‌فرض (تعدادِ هسته − ۱) نخ دارد و بین کلِ JVM مشترک است — همان استخری که parallelStream() هم استفاده می‌کند. اگر داخلِ یک مرحله I/O بلاک کنی، نخ‌های commonPool را می‌خوری و ناگهان parallelStream()های بی‌ربطِ جای دیگرِ برنامه کند می‌شوند. قانون: هیچ‌وقت کارِ بلاک‌کننده یا I/O را روی commonPool اجرا نکن؛ همیشه یک Executor صریح به *Async بده.

مدیریتِ استثنا هم تلهٔ خودش را دارد: در یک زنجیره، یک استثنا مراحلِ بعدی را رد می‌کند تا به exceptionally/handle برسد.

fetchUser(id)
    .thenApply(this::risky)
    .exceptionally(ex -> User.GUEST)      // فقط مسیرِ خطا را می‌گیرد و مقدار می‌دهد
    .thenAccept(this::render);
// handle(value, ex) هر دو مسیر را می‌گیرد؛ whenComplete side-effect است و استثنا را تغییر نمی‌دهد

از Java 9 هم orTimeout(...) و completeOnTimeout(...) اضافه شد تا دیگر مجبور نباشی timeout را دستی بسازی.

الگوی Memoizer: شکستِ cache stampede

فصل با computeIfAbsent مسابقهٔ get-then-put را حل کرد. اما یک مشکلِ عمیق‌تر هست: اگر محاسبه گران و کند باشد و ۱۰۰ نخ هم‌زمان همان کلیدِ غایب را بخواهند، آیا ۱۰۰ بار محاسبه می‌شود؟ راه‌حلِ کلاسیکِ Goetz این است که به‌جای مقدار، خودِ CompletableFuture را در مپ کش کنی:

ConcurrentHashMap<K, CompletableFuture<V>> cache = new ConcurrentHashMap<>();
V get(K key) {
    CompletableFuture<V> f = cache.computeIfAbsent(key, k -> CompletableFuture.supplyAsync(() -> compute(k)));
    return f.join();
}

حالا اولین نخ فیوچر را می‌سازد و بقیه همان فیوچرِ در حالِ اجرا را می‌گیرند و منتظرش می‌مانند — محاسبه فقط یک بار انجام می‌شود. این «thundering herd / cache stampede» را می‌کُشد. (نکته: اگر محاسبه شکست خورد، فیوچرِ خرابِ کش‌شده را remove کن تا دوباره تلاش ممکن شود.)

double-checked locking و اصطلاحِ holder

یک الگوی کلاسیک که فصل نگفت: مقداردهیِ تنبل و thread-safe بدونِ قفل‌گرفتن در مسیرِ داغ. نسخهٔ ساده‌لوحانه‌اش سال‌ها شکسته بود:

private Helper helper; // BUG: بدونِ volatile
Helper get() {
    if (helper == null) {                 // چک ۱ (بدونِ قفل)
        synchronized (this) {
            if (helper == null)           // چک ۲ (با قفل)
                helper = new Helper();
        }
    }
    return helper;
}
چرا `volatile` اینجا حیاتی است

helper = new Helper() یک عملیاتِ اتمی نیست: تخصیصِ حافظه، اجرای سازنده، و انتسابِ رفرنس. بدونِ volatile، JMM اجازه می‌دهد این‌ها بازچینش شوند — یعنی رفرنس می‌تواند قبل از پایانِ سازنده منتشر شود. نخِ دومی که چکِ اولِ بدونِ قفل را رد می‌کند، یک رفرنسِ غیر-null اما به یک ابجکتِ نیمه‌ساخته می‌گیرد. راه‌حل: private volatile Helper helper; — که یک لبهٔ happens-before می‌سازد و بازچینش را ممنوع می‌کند.

اما راهِ بهتر و تمیزتری هست که اصلاً به volatile و double-check نیاز ندارد — اصطلاحِ initialization-on-demand holder:

class Config {
    private Config() { /* گران */ }
    private static class Holder { static final Config INSTANCE = new Config(); }
    static Config get() { return Holder.INSTANCE; } // تنبل + thread-safe، رایگان
}
چرا holder از هر قفلی بهتر است

JVM تضمین می‌کند که یک کلاس فقط یک بار و به‌صورت thread-safe مقداردهی می‌شود (زیر یک قفلِ داخلیِ کلاس). کلاسِ Holder تا اولین ارجاع به Holder.INSTANCE بارگذاری نمی‌شود — پس مقداردهی تنبل است — و انحصارِ متقابل را خودِ classloader مجانی به تو می‌دهد. نه volatile، نه synchronized، نه هیچ هزینه‌ای در مسیرِ داغ. برای singletonها این ایده‌آل است (یا یک enum تک‌عضوی).

لایهٔ سخت‌افزار: false sharing

بعضی باگ‌های کارایی هیچ ربطی به منطقِ کد ندارند و در خطِ کش (cache line) پردازنده زندگی می‌کنند. پردازنده حافظه را نه بایت‌به‌بایت، بلکه در بلوک‌های ۶۴‑بایتی (خطِ کش) جابه‌جا می‌کند.

دو نفر و یک وایت‌بردِ مشترک

دو کارگر روی دو گوشهٔ مختلفِ یک وایت‌برد می‌نویسند. منطقاً به هم کاری ندارند. اما وایت‌برد آن‌قدر کوچک است که هر بار یکی می‌نویسد، سیستم مجبور است کلِ برد را برای دیگری «بی‌اعتبار» و دوباره کپی کند. آن‌ها داده‌ای به اشتراک نمی‌گذارند اما مکانِ فیزیکی را به اشتراک می‌گذارند — و همین کند می‌کندشان.

اگر دو متغیرِ مستقل که نخ‌های مختلف روی‌شان می‌نویسند، اتفاقاً در یک خطِ کش بیفتند، هر نوشتنِ یکی، کشِ آن یکی را invalidate می‌کند و پروتکلِ coherency پردازنده مدام آن خط را بین هسته‌ها پینگ‌پونگ می‌کند. کد درست است اما شاید ۵ تا ۱۰ برابر کند. اسمش false sharing است چون واقعاً اشتراکی نیست.

چطور در عمل می‌بینی و درمانش می‌کنی

اگر یک شمارندهٔ داغ داری و مشکوکی، جاوا @jdk.internal.vm.annotation.Contended (یا نسخهٔ عمومی‌اش) دارد که فیلد را با padding از بقیه جدا می‌کند — و باید JVM را با -XX:-RestrictContended اجرا کنی. اما راهِ سالم‌تر معمولاً استفاده از ابزارهایی است که خودشان padding دارند: دقیقاً به همین دلیل LongAdder سریع‌تر از یک آرایه از AtomicLong است — سلول‌هایش padding‌شده‌اند. این یک سؤالِ مصاحبهٔ خوب برای نقش‌های low-latency (fintech, trading) است.

CAS و مسئلهٔ ABA

فصل از AtomicLong و CAS گفت اما از یکی از ظریف‌ترین باگ‌های الگوریتم‌های non-blocking نگفت: مسئلهٔ ABA. CAS می‌گوید «اگر مقدار هنوز A است، به B تغییرش بده». اما اگر بین خواندنِ تو و CASِ تو، مقدار از A به X و دوباره به A برگشته باشد چه؟ CASِ تو موفق می‌شود، چون فقط مقدار را می‌بیند نه تاریخچه را.

کلیدِ خانه‌ای که عوض و دوباره نصب شده

از خانه بیرون می‌روی، در را با کلیدت چک می‌کنی: قفلِ «A». برمی‌گردی، باز قفلِ «A» است، پس فکر می‌کنی «هیچ‌چیز عوض نشده». اما در این فاصله کسی کلِ قفل را باز کرده، خانه را خالی کرده و یک قفلِ ظاهراً یکسان دوباره نصب کرده. مقدار یکی است، اما دنیا زیرِ پایت عوض شده.

در ساختارهای لینک‌شدهٔ lock-free (مثل یک stackِ Treiber که گره‌ها را از یک free-list دوباره استفاده می‌کند)، ABA می‌تواند یک گرهِ آزادشده را دوباره وصل کند و ساختار را خراب کند. راه‌حل: به هر مقدار یک شماره نسخه (stamp) بچسبان.

AtomicStampedReference<Node> top = new AtomicStampedReference<>(head, 0);
int[] stampHolder = new int[1];
Node cur = top.get(stampHolder);
// ... CAS با مقدار *و* نسخهٔ بعدی؛ حتی اگر مقدار به cur برگردد، stamp فرق می‌کند
top.compareAndSet(cur, next, stampHolder[0], stampHolder[0] + 1);

AtomicStampedReference مقدار و یک شمارندهٔ int را با هم اتمیک می‌کند؛ حالا A→X→A دیگر گول نمی‌زند چون stamp پیش رفته. (AtomicMarkableReference نسخهٔ boolean‌ی برای علامت‌گذاریِ منطقی حذف است.)

نقشهٔ همزمانیِ جاوای مدرن (۲۰۲۵–۲۰۲۶)

چند تغییرِ بزرگ که هر سنیوری باید در مصاحبهٔ ۲۰۲۶ بداند:

biased locking مُرد

تا مدت‌ها JVM یک بهینه‌سازی به‌نامِ biased locking داشت (فرضِ اینکه یک قفل معمولاً همیشه دستِ همان نخ است). این در JDK 15 با JEP 374 پیش‌فرض غیرفعال و بعد کاملاً حذف شد، چون با کدِ به‌شدت concurrentِ امروزی بیشتر ضرر داشت تا فایده. نتیجهٔ عملی: synchronizedِ بدون‌رقابت امروز کمی گران‌تر از قدیم است — دلیلی بیشتر برای نگه‌داشتنِ ناحیه‌های بحرانیِ کوچک.

pinning در JDK 24 حل شد (JEP 491)

فصل درست گفت که در Java 21 اگر یک virtual thread داخلِ synchronized بلاک شود، به carrier thread pin (میخکوب) می‌شود و مقیاس‌پذیری را می‌شکند. خبرِ بزرگ: JEP 491 در JDK 24 این را تقریباً کامل حل کرد — حالا مانیتور به خودِ virtual thread وصل است، پس نخ می‌تواند حتی داخلِ synchronized بلاک شود و carrier را آزاد کند. یعنی توصیهٔ «برای virtual threadها ReentrantLock را به synchronized ترجیح بده» عمدتاً به تاریخِ پیش از JDK 24 مربوط است. (pinning هنوز در فریم‌های native و class initializerها باقی است، اما آن‌ها نادرند.)

Scoped Values جانشینِ ThreadLocal شد (JDK 25، JEP 506)

فصل درست از خطرِ نشتِ ThreadLocal روی استخرها گفت. جاوای مدرن جایگزینِ بهتری دارد: Scoped Values که در JDK 25 نهایی (final) شد. یک مقدارِ تغییرناپذیر را برای طولِ یک عملیات و همهٔ زیرکارهایش (و نخ‌های فرزندش) به اشتراک می‌گذاری و در پایانِ scope خودکار پاک می‌شود — نه remove()ِ دستی، نه نشت، و ارزان‌تر از ThreadLocal مخصوصاً با میلیون‌ها virtual thread:

private static final ScopedValue<User> CURRENT = ScopedValue.newInstance();
ScopedValue.where(CURRENT, user).run(() -> handleRequest()); // در scope قابل‌خواندن، بیرونش نه

همراهِ آن، Structured Concurrency (StructuredTaskScope) هست که هنوز preview است (پنجمین preview در JDK 25) و می‌گذارد گروهی از زیرکارها را مثلِ یک واحد مدیریت کنی: یا همه موفق، یا همه با هم لغو — پایانِ نخ‌های سرگردانِ leak‌شده.

سؤالاتِ سختِ مصاحبهٔ سنیور

۱) یک `ThreadPoolExecutor` با core=5, max=50 و `LinkedBlockingQueue` بی‌ظرفیت ساخته‌ای. زیرِ بارِ سنگین چند نخ خواهی داشت؟

دقیقاً ۵. با صفِ نامحدود، offer هیچ‌وقت شکست نمی‌خورد، پس منطقِ استخر هرگز به مرحلهٔ ساختِ نخِ اضافه (تا max) نمی‌رسد؛ maximumPoolSize کاملاً بی‌اثر است. تسک‌ها بی‌صدا در صف تلنبار می‌شوند تا OOM. درست: از یک صفِ محدود (ArrayBlockingQueue) به‌همراه یک RejectedExecutionHandler مثلِ CallerRunsPolicy استفاده کن تا هم از max بهره ببری و هم backpressure واقعی داشته باشی.

۲) (تله) این کد چه اشکالی دارد؟
try { doBlockingWork(); }
catch (InterruptedException e) { log.warn("interrupted"); }

سیگنالِ لغو را بلعیده. وقتی InterruptedException پرتاب شد، JVM پرچمِ interrupt را پاک کرده؛ این کد آن را restore نمی‌کند و به بالا هم propagate نمی‌کند. نتیجه: لایه‌های بالاتر (و shutdownNow()/timeout) دیگر نمی‌فهمند این نخ باید متوقف شود و نخ برای همیشه زنده می‌ماند. درست: یا استثنا را throw کن، یا Thread.currentThread().interrupt() را صدا بزن و از تسک خارج شو.

۳) تفاوتِ `thenApply` و `thenCompose` در `CompletableFuture` چیست و چرا مهم است؟

thenApply مثلِ map است: یک تابعِ همگام T → U می‌گیرد. thenCompose مثلِ flatMap است: یک تابعِ T → CompletableFuture<U> می‌گیرد و نتیجه را صاف می‌کند. اگر با thenApply تابعی بدهی که خودش فیوچر برمی‌گرداند، به CompletableFuture<CompletableFuture<U>>ِ تودرتو می‌رسی که کار با آن کابوس است. هر جا مرحلهٔ بعدی خودش ناهمگام است (یک فراخوانِ سرویسِ دیگر)، thenCompose درست است.

۴) (سخت) چرا اجرای کارِ بلاک‌کننده روی `CompletableFuture.supplyAsync(...)` بدونِ Executor خطرناک است؟

چون پیش‌فرض روی ForkJoinPool.commonPool() می‌رود که (الف) فقط هسته−۱ نخ دارد و (ب) بین کلِ JVM مشترک است — همان استخرِ parallelStream(). اگر داخلش I/O بلاک کنی، نخ‌های محدودِ commonPool را اشغال می‌کنی و کارهای بی‌ربطِ موازیِ جای دیگرِ برنامه گرسنه می‌شوند؛ در بدترین حالت با کارهای بازگشتیِ fork/join یک self-deadlock. درست: همیشه یک Executor صریحِ اختصاصی (یا یک virtual-thread executor) به نسخهٔ *Async بده.

۵) (تله) double-checked locking بدونِ `volatile` چطور می‌شکند؟

instance = new Helper() سه گام است: تخصیص، اجرای سازنده، انتساب. بدونِ volatile، JMM اجازهٔ بازچینش می‌دهد؛ پس رفرنس ممکن است قبل از پایانِ سازنده منتشر شود. نخِ دومی که چکِ اولِ بدونِ‌قفل را می‌بیند، یک رفرنسِ غیر-null اما به ابجکتِ نیمه‌ساخته می‌گیرد و از آن استفاده می‌کند. فیلد باید volatile باشد تا لبهٔ happens-before بسازد. راهِ بهتر: اصطلاحِ initialization-on-demand holder که کلاً از این تله فرار می‌کند.

۶) (سخت) false sharing چیست و چرا کدِ کاملاً درست را کند می‌کند؟

پردازنده حافظه را در خطوطِ کشِ ۶۴‑بایتی جابه‌جا می‌کند. اگر دو متغیرِ مستقل که نخ‌های مختلف رویشان می‌نویسند در یک خطِ کش بیفتند، هر نوشتنِ یکی نسخهٔ کشِ آن یکی را invalidate می‌کند و پروتکلِ cache coherency آن خط را مدام بین هسته‌ها پینگ‌پونگ می‌کند. منطقاً هیچ اشتراکی نیست، اما کارایی چند برابر افت می‌کند. درمان: padding (مثلِ @Contended) یا استفاده از ساختارهای padding‌شده مثلِ LongAdder. مثالِ عالی برای این‌که «درستی ≠ کارایی».

۷) مسئلهٔ ABA در CAS را توضیح بده و چطور حلش می‌کنی.

CAS فقط مقدار را می‌بیند، نه تاریخچه را. اگر مقدار از A به B و دوباره به A برگردد، CASِ تو موفق می‌شود انگار هیچ‌چیز عوض نشده — در حالی که ممکن است invariantی که به آن تکیه داشتی نقض شده باشد (مثلاً یک گرهٔ آزاد و دوباره‌استفاده‌شده در یک stackِ lock-free). حل: به هر مقدار یک شمارهٔ نسخه بچسبان با AtomicStampedReference، که مقدار و stamp را با هم اتمیک مقایسه می‌کند؛ چون stamp همیشه پیش می‌رود، بازگشتِ A→X→A دیگر گول نمی‌زند.

۸) در جاوای ۲۰۲۵/۲۰۲۶، توصیهٔ «برای virtual threadها به‌جای `synchronized` از `ReentrantLock` استفاده کن» هنوز درست است؟

عمدتاً یک توصیهٔ پیش از JDK 24 است. در Java 21، بلاک‌شدنِ یک virtual thread داخلِ synchronized آن را به carrier thread میخکوب (pin) می‌کرد و مقیاس‌پذیری را می‌شکست. اما JEP 491 در JDK 24 این را حل کرد: مانیتور حالا به خودِ virtual thread وصل است و نخ می‌تواند داخلِ synchronized هم unmount شود. پس روی JDK 24+، synchronized دیگر pin نمی‌کند (به‌جز موارد نادرِ فریمِ native و class initializer). این نشان می‌دهد که «best practice»ها نسخه‌دار‌ند و باید بدانی برای کدام JDK حرف می‌زنی.

جمع‌بندیِ نکاتِ سنیور

پشتِ هر سرویس یک ThreadPoolExecutor است: صفِ نامحدود maximumPoolSize را به دروغی تزئینی بدل می‌کند — صفِ محدود + CallerRunsPolicy بزن. interruption یک پروتکلِ همکاری است، نه کلیدِ کشتن؛ هیچ‌وقت InterruptedException را نبلع — یا throw کن یا پرچم را restore کن. در CompletableFuture، thenCompose همان flatMap است، و هیچ‌وقت روی commonPool بلاک نکن؛ الگوی Memoizer با کش‌کردنِ خودِ فیوچر، cache stampede را می‌کُشد. برای مقداردهیِ تنبل، holder idiom را به double-checked locking (که بدونِ volatile می‌شکند) ترجیح بده. بعضی افت‌های کارایی در سطحِ false sharing خطِ کش‌اند، نه منطق. CAS در برابرِ ABA آسیب‌پذیر است — با AtomicStampedReference نسخه‌دارش کن. و نقشهٔ مدرن را بدان: biased locking مرد، pinning در JDK 24 حل شد، و Scoped Values جانشینِ امنِ ThreadLocal است.

جمع‌بندی

هر باگ همزمانی یا ایمنی را می‌شکند (پاسخ غلط: رقابت، بررسی-سپس-عمل) یا زندگی‌مندی را (معلق‌شدن: بن‌بست، زنده‌قفلی، گرسنگی) — و راه‌حل‌هایشان در جهت مخالف‌اند، پس روی آن لبهٔ باریک تعادل نگه دار. زیربنای همه‌چیز، رابطهٔ happens-before در JMM است؛ بدون آن هیچ تضمینی نداری. بن‌بست به هر چهار شرط کافمن نیاز دارد — فقط یکی را بشکن، معمولاً با ترتیب سراسری قفل یا tryLock+مهلت (و آن finally حیاتی). زنده‌قفلی را با عدم تقارن (عقب‌نشینی تصادفی)، گرسنگی را با قفل‌های عادل درمان کن. برای رقابت‌ها، از عملیات اتمیِ واحد استفاده کن: computeIfAbsent، AtomicLong/LongAdder. برای تولیدکننده-مصرف‌کننده، BlockingQueue کران‌دار (فشار برگشتی) و قرص سمّی؛ در نسخهٔ دستی همیشه while نه if و دو Condition. ابزارهای هماهنگی را بشناس (latch یک‌بارمصرف، barrier تکرارپذیر و شکننده، semaphore با خطر release اضافه، phaser پویا). و مهم‌تر از همه: ارزان‌ترین همزمانی، تغییرناپذیری و محصورسازی است — اگر چیزی مشترکِ تغییرپذیر نباشد، هیچ باگی هم ندارد. درستی را استدلال می‌کنی، آزمون فقط پس‌رفت‌ها را می‌گیرد.

Let's be honest with each other: concurrency is where even strong programmers get humbled. Code that runs flawlessly on a single thread can suddenly produce wrong answers or hang forever the moment two threads touch the same data. The good news is that all of these disasters fall into a small number of families, and each family has a battle-tested pattern that defeats it. In this lesson we'll build that whole map in your head — from scratch.

Roadmap for this lesson

First we build one big compass: every concurrency bug either violates safety or liveness. Then we meet four monsters: deadlock, livelock, starvation, and race conditions. Next we learn the defensive tools one by one: BlockingQueue for producer-consumer, the coordination primitives (latch/barrier/semaphore/phaser), and finally two structural strategies that eliminate whole bug classes at the root — immutability and confinement. The lesson closes with a full interview-questions section.

Part 0 — words you must know

Before anything else, let's unpack a few words that recur throughout, each with an analogy, so you never get lost later.

  • Thread: an independent line of execution. Picture each thread as a worker busy at the same time as the others.
  • Lock: the key to a room. Only one worker can hold the key at a time; the rest wait outside. In Java, synchronized and the ReentrantLock class create this key.
  • Reentrant: a worker who holds a room's key can re-enter the same room without blocking. Java's synchronized monitors are reentrant.
  • Atomic: an operation that either happens completely or not at all — no other worker can observe a half-finished state mid-operation.
  • Interleaving: the order in which the OS weaves together the steps of different workers. You have no control over this order, and that's exactly where bugs are born.
A busy restaurant kitchen

You can understand this entire lesson through the image of a busy kitchen: threads are the cooks, locks are the shared tools (one knife, one stove), and shared data is the food on the counter. When two cooks grab the same pot without coordinating, either the food gets ruined (a safety violation) or they both wait for each other and no food ever comes out (a liveness violation).

Mental model: liveness vs. safety

Let's start with the biggest idea, because the rest of the lesson lives under its shadow. Every concurrency bug, without exception, violates one of two properties:

  • Safety means "nothing bad ever happens." A race condition corrupts state, a HashMap resizes into an infinite loop, a check-then-act sees a stale value. Safety failures produce wrong answers.
  • Liveness means "something good eventually happens." Deadlock, livelock, and starvation mean threads stop making progress. Liveness failures produce hangs.

Why does this split matter so much? Because the fixes are opposite in spirit, and if you don't grasp that, fixing one bug spawns another.

Two kinds of kitchen failure

A safety failure is like two cooks both salting the same dish so it comes out inedibly salty — the food was produced, but it's wrong. A liveness failure is like two cooks each waiting for the other to release the stove first — the food never gets made, even though neither made an obvious mistake. One is a "wrong result," the other is "no result at all."

Senior engineers internalize this split because the fixes conflict. Safety usually needs more coordination (locks, atomics, happens-before). Liveness usually needs less or smarter coordination (lock ordering, timeouts, fairness, non-blocking algorithms). Over-synchronizing to fix a race can create a deadlock; loosening locks to fix a deadlock can reintroduce a race.

You are always balancing on a narrow ridge

On one side is the race (needs more locking), on the other is deadlock (needs less or smarter locking). Good engineering isn't falling to either side — it's staying on that ridge. Every time you're about to add a lock, ask yourself: "does this create a new deadlock?"

The invisible foundation: the Java Memory Model (JMM)

Now a deeper layer. The Java Memory Model (JMM) underpins all of it. Its key term is happens-before: a guaranteed "edge," or relationship, between two operations that says the first is definitely visible to the second.

A note on a whiteboard

Imagine worker A writes something in their private notebook. Until A publishes it onto the shared whiteboard, worker B might see an old or half-written version — or nothing at all. The happens-before relationship is that "publishing to the whiteboard": the guarantee that what A wrote reaches B correctly and completely.

Without a happens-before edge between a write on thread A and a read on thread B, B may see a stale value, a partially constructed object, or reordered operations — even on x86. What establishes these edges? Locks (synchronized, ReentrantLock), volatile, final fields (after construction), Thread.start/join, and the java.util.concurrent classes.

"It worked on my machine" is not an argument

If you can't explain which happens-before relationship guarantees thread B sees thread A's write, your code merely got lucky — and luck runs out on different hardware or under heavy load. Always have a happens-before argument.


Deadlock: the four Coffman conditions

An intersection with no traffic lights

Picture four cars arriving from four directions at an intersection with no lights, each wanting to enter but each waiting for the car on its right to go first. Nobody moves, because everybody is waiting on someone else. That's exactly deadlock: a wait cycle that never opens.

A deadlock is a cycle of threads each holding a resource the next one wants. The wonderfully practical fact is that a deadlock requires all four Coffman conditions simultaneously — and if you break just one, deadlock becomes completely impossible.

Condition Meaning How to break it
Mutual exclusion A resource is held exclusively Use immutable/shared-read resources, lock-free structures
Hold and wait A thread holds one lock while requesting another Acquire all locks at once, or release before requesting
No preemption Locks can't be forcibly taken Use tryLock with timeout + backoff
Circular wait A cycle exists in the wait-for graph Impose a global lock ordering

Let's see the most classic deadlock bug: two bank accounts and two threads transferring money in opposite directions.

// BUG: lock order depends on argument order → circular wait
void transfer(Account from, Account to, long amount) {
    synchronized (from) {
        synchronized (to) {          // T1: A then B; T2: B then A → deadlock
            from.debit(amount);
            to.credit(amount);
        }
    }
}

What happens? Thread 1 calls transfer(A, B, ...) and grabs lock A. At that same instant thread 2 calls transfer(B, A, ...) and grabs lock B. Now thread 1 waits for B (held by thread 2) and thread 2 waits for A (held by thread 1). The wait cycle is complete; both block forever.

Argument order decides lock order — and that's the disaster

The root of the bug is that lock-acquisition order is tied to whatever order the caller passed the arguments. As long as two calls with reversed argument order exist, circular wait is lurking. The fix must sever that dependency.

Fix 1: global lock ordering (breaks circular wait)

The idea is simple and elegant: give every lockable object a stable, unique ordering key and always acquire in that order, regardless of the order the caller passed.

void transfer(Account from, Account to, long amount) {
    Account first  = from.id() < to.id() ? from : to;  // total order by id
    Account second = from.id() < to.id() ? to : from;
    synchronized (first) {
        synchronized (second) {
            from.debit(amount);
            to.credit(amount);
        }
    }
}
Why this works

If everyone always locks the smaller-id object first, it becomes impossible for two threads to lock in opposite order. The cycle that deadlock requires can never form. This "global ordering" is the single most practical anti-deadlock weapon in most real systems.

Now an edge case: if from.id() == to.id() (a self-transfer) you'd synchronized the same monitor twice — harmless because Java monitors are reentrant (remember? a worker holding the key can re-enter), but you should still guard against a same-account transfer semantically. And when no natural unique key exists, use System.identityHashCode and a tie-breaker lock for the rare collision of equal hashes:

private static final Object TIE = new Object();

void transfer(Account from, Account to, long amount) {
    int hf = System.identityHashCode(from), ht = System.identityHashCode(to);
    if (hf < ht)       lockedTransfer(from, to, amount);
    else if (hf > ht)  lockedTransfer(to, from, amount, /*reverse*/ true);
    else synchronized (TIE) { lockedTransfer(from, to, amount); } // hash collision
}

The TIE lock covers the rare case where two distinct objects happen to have equal identityHashCodes and can't be ordered; in that one rare case a single shared lock keeps the ordering consistent.

Fix 2: tryLock with timeout (breaks no preemption)

The second strategy, instead of ordering locks, lets a thread escape an eternal wait: rather than "wait however long it takes," say "try for at most 50 milliseconds; if it fails, give up and retry."

boolean transfer(Account from, Account to, long amount, Duration timeout)
        throws InterruptedException {
    long deadline = System.nanoTime() + timeout.toNanos();
    while (System.nanoTime() < deadline) {
        if (from.lock.tryLock(50, TimeUnit.MILLISECONDS)) {
            try {
                if (to.lock.tryLock(50, TimeUnit.MILLISECONDS)) {
                    try { from.debit(amount); to.credit(amount); return true; }
                    finally { to.lock.unlock(); }
                }
            } finally { from.lock.unlock(); } // ALWAYS release the outer lock
        }
        // couldn't get both: back off a randomized amount to avoid livelock
        Thread.sleep(ThreadLocalRandom.current().nextInt(1, 10));
    }
    return false;
}
The critical detail is that outer `finally`

If you fail to get to but keep holding from while you retry, you've just created "hold and wait" — the very deadlock you were trying to escape! That finally releasing from turns a transient failure into a clean backoff instead of an eternal lock. And the randomized backoff (not a fixed one) is what stops the retry loop from becoming a livelock.

Detecting deadlock in production

Suppose that despite all precautions, a server hangs in production. How do you confirm it's a deadlock?

  • Thread dump: jstack <pid> (or kill -3) prints a "Found one Java-level deadlock" section with the exact cycle. This is your first move on any hung JVM.
  • Programmatic: ThreadMXBean.findDeadlockedThreads() can run on a watchdog thread and alert automatically.
ThreadMXBean mx = ManagementFactory.getThreadMXBean();
long[] deadlocked = mx.findDeadlockedThreads(); // null if none
if (deadlocked != null) log.error("DEADLOCK: {}", Arrays.toString(deadlocked));

Livelock and starvation

Deadlock isn't the only way to hang. It has two close cousins that are more deceptive because the threads appear busy.

Two people in a corridor

Two people meet face to face in a narrow corridor. Both politely step to one side — still facing each other. Both step the other way — still facing each other. They are not blocked; they are actively moving, yet they never get past. That's livelock.

Livelock: threads are not blocked — they're actively running — but keep responding to each other and make no progress. In code it appears when threads back off and retry in lockstep, or when message-passing actors keep handing a task back and forth. The cure is asymmetry: randomized backoff (the nextInt(1, 10) above), or a priority/token that breaks the symmetry. If both people in the corridor flipped a coin to decide who goes first, the symmetry breaks and the problem is solved.

A bakery with no orderly queue

Starvation is like a bakery with no proper line, where each time whoever pushes hardest gets served. A polite, quiet person might stand for hours and never get served — not because they're locked, but because others keep cutting ahead.

Starvation: a thread never gets a resource because others perpetually win the race. Causes include unfair locks under heavy contention, thread-priority abuse, and a busy writeLock starving readers (or vice versa). Fixes:

  • Use fair locks where latency tails matter: new ReentrantLock(true). A fair lock is like the orderly bakery queue — first come, first served. Fairness trades throughput for bounded waiting, so measure before defaulting to it.
  • Prefer ReentrantReadWriteLock with a fairness/downgrade policy, or StampedLock for read-heavy workloads. Note: StampedLock is not reentrant and its optimistic reads must be validated (we'll see it shortly).

Race conditions and check-then-act

A race condition is when a program's correctness depends on the interleaving of threads — an order you don't control.

Two people and an empty fridge

You open the fridge, see the milk is gone, and head out to buy some. Your housemate saw the same empty fridge at that same moment and also headed out. Result: two people buy milk. You both "checked" (no milk) and then "acted" (bought), but between check and act the state shifted under you. This is exactly the check-then-act pattern.

The most common shape of a race is check-then-act: you observe a value and act on it, but the value changes in between.

// BUG: classic check-then-act — two threads can both pass the null check
private Connection conn;
Connection get() {
    if (conn == null) {          // check
        conn = open();           // act — two connections leak, or worse
    }
    return conn;
}

ConcurrentHashMap invites a subtler version of the same trap — subtle because the map itself is thread-safe, so people assume the danger is gone:

// BUG: get-then-put is not atomic; two threads compute twice, one wins
Value v = map.get(key);
if (v == null) {
    v = expensiveCompute(key);
    map.put(key, v);             // last write wins; wasted work; inconsistent v
}

The point is this: each ConcurrentHashMap operation is atomic on its own, but when you place a get and a put side by side, that combination is no longer atomic. The fix is to use a single atomic operation — which is precisely why the concurrent collections exist:

// computeIfAbsent runs the mapping function atomically per key
Value v = map.computeIfAbsent(key, this::expensiveCompute);
Real gotcha: recursive `computeIfAbsent`

In Java 8, calling computeIfAbsent recursively on the same map for a different key, inside the mapping function, can corrupt the table or deadlock. Java 9+ detects this reentrant modification and throws. Simple rule: never do map-mutating work on that same map inside the mapping function.

Now a small but ubiquitous race — the compound "read-modify-write" on a counter:

count++;                         // BUG: read, add, write — three steps, not atomic

That innocent ++ is actually three steps: read the value, add one, write it back. Two threads can read the old value at the same time and both write old+1 — one increment is lost. Fixes, in ascending order of scalability:

synchronized (lock) { count++; }             // correct, contended
AtomicLong count = ...; count.incrementAndGet(); // CAS, better under moderate contention
LongAdder adder = ...; adder.increment();        // striped, best under HIGH contention
Why `LongAdder` wins under heavy contention

AtomicLong has a single memory location everyone fights over — like one teller window with a long queue behind it. LongAdder spreads the work across several cells — like opening several windows — and only sums them when you call sum(). So when many threads write and you read rarely, LongAdder wins. But if you read the value constantly or contention is low, AtomicLong is simpler and plenty.


Producer–consumer with BlockingQueue

The pass window between kitchen and dining room

Picture a limited pass window between the cooks (producers) and the waiters (consumers). The cook places a plate on the window; a waiter picks it up. If the window fills, the cook must wait (rather than piling plates on the floor); if it's empty, the waiter waits. That automatic two-way waiting is exactly what a BlockingQueue does for you.

Hand-rolling wait/notify producer-consumer is a rite of passage and a source of endless bugs (missed signals, lost wakeups, notify vs notifyAll). In production you almost always use a BlockingQueue, which encapsulates the bounded buffer, the condition waiting, and the backpressure all in one.

BlockingQueue<Task> queue = new ArrayBlockingQueue<>(1000); // bounded → backpressure

// Producer
void produce(Task t) throws InterruptedException {
    queue.put(t);   // BLOCKS when full — this is desirable backpressure
}

// Consumer with a poison-pill shutdown
static final Task POISON = new Task.Poison();
void consumeLoop() throws InterruptedException {
    while (true) {
        Task t = queue.take();        // BLOCKS when empty
        if (t == POISON) { queue.put(POISON); return; } // re-insert for siblings
        handle(t);
    }
}

"Backpressure" is an important term worth unpacking right here: it means that when the consumer is slow, that slowness propagates backward — to the producer — and calms it down too. Without backpressure, a fast producer fills memory until the program crashes. Key choices:

  • Bounded (ArrayBlockingQueue, bounded LinkedBlockingQueue) gives backpressure — producers slow down instead of the heap exploding. Prefer bounded. An unbounded queue turns a load spike into an OutOfMemoryError.
  • SynchronousQueue has zero capacity: every put hands directly to a take — like passing a hot plate hand-to-hand rather than setting it on the counter. It's the engine behind Executors.newCachedThreadPool and forces true rendezvous.
  • Poison pill is the clean shutdown idiom: enqueue a sentinel so consumers exit after draining the real work, rather than being interrupted mid-work. Re-insert it so multiple consumers all see it.

The hand-rolled version (know it for interviews)

Even though you'll use BlockingQueue in practice, interviewers love to see you build a bounded buffer by hand. This is the correct version:

// Correct bounded buffer with a single lock and two conditions
class BoundedBuffer<E> {
    private final Object[] buf;
    private int count, head, tail;
    private final ReentrantLock lock = new ReentrantLock();
    private final Condition notFull  = lock.newCondition();
    private final Condition notEmpty = lock.newCondition();

    BoundedBuffer(int cap) { buf = new Object[cap]; }

    void put(E e) throws InterruptedException {
        lock.lock();
        try {
            while (count == buf.length) notFull.await(); // while, NOT if
            buf[tail] = e; tail = (tail + 1) % buf.length; count++;
            notEmpty.signal();
        } finally { lock.unlock(); }
    }
    @SuppressWarnings("unchecked")
    E take() throws InterruptedException {
        lock.lock();
        try {
            while (count == 0) notEmpty.await();
            E e = (E) buf[head]; buf[head] = null;       // null out for GC
            head = (head + 1) % buf.length; count--;
            notFull.signal();
            return e;
        } finally { lock.unlock(); }
    }
}

Two senior-level rules are baked into this code. First: always wait in a while, never an if. Why? Because between the moment you're signaled and the moment you reacquire the lock, another thread may have snatched that empty slot; and per the Java spec, a "spurious wakeup" is also allowed — meaning you can sometimes wake with no signal at all. The only safe pattern is re-checking the predicate in a loop. Second: use two separate conditions (notFull and notEmpty) so a signal on "not full" never wastefully wakes a consumer waiting on "not empty."

`if` instead of `while` = an unreproducible corruption

If you use if instead of while, a thread that woke spuriously or late proceeds without rechecking and acts on the wrong state — e.g., taking from a buffer that just refilled. This kind of bug only strikes occasionally, under specific load, and is nearly impossible to reproduce. With a single monitor you'd be forced to notifyAll, which is O(number of waiters) wasteful.


Coordination primitives: latches, barriers, semaphores, phasers

So far we worked with locks. But sometimes you don't want "exclusive access" — you want to coordinate threads, like "everyone start together" or "everyone wait until the last one arrives." Java ships ready-made tools for this.

Coordination tools as track-meet gear

CountDownLatch is like the starting pistol — fired once, everyone runs. CyclicBarrier is like a line all runners gather at after each lap, then start the next lap together, over and over. Semaphore is like a limited number of pool lanes — only N swimmers at once. Phaser is like a flexible coach who can add or remove runners mid-workout.

Primitive Reusable? Use case
CountDownLatch No (one-shot) Wait for N events to complete before proceeding
CyclicBarrier Yes N threads meet at a barrier repeatedly (phased computation)
Semaphore Yes Limit concurrent access to N permits (pool, rate cap)
Phaser Yes Dynamic party count, multi-phase; flexible barrier
Exchanger Yes Two threads swap objects at a rendezvous

Let's see CountDownLatch in action — the classic "start everyone together, then wait for everyone to finish" pattern:

// CountDownLatch: start N workers together, wait for all to finish
CountDownLatch ready = new CountDownLatch(1);   // release gate
CountDownLatch done  = new CountDownLatch(N);
for (int i = 0; i < N; i++) new Thread(() -> {
    ready.await();          // all block until the gate opens
    work();
    done.countDown();       // signal completion
}).start();
ready.countDown();          // fire the starting gun
done.await();               // main waits for everyone

Here ready is a gate that opens when its count hits zero; all workers wait behind it until main fires the pistol. done starts at N, and each worker counts it down as it finishes; when it hits zero, main is released.

CountDownLatch counts down to zero and stays there — it cannot be reset, it's one-shot. When you need a repeatable rendezvous, use CyclicBarrier, which can run a barrier action when the last thread arrives and then resets itself:

CyclicBarrier barrier = new CyclicBarrier(N, () -> mergePhaseResults());
// each worker calls barrier.await() at the end of every phase
`CyclicBarrier` is all-or-nothing

If one of the waiting threads is interrupted or times out, the barrier is broken and every other waiter gets a BrokenBarrierException — not just that one thread. It's like a roped-together climbing team: if one falls, they drag everyone. Robust code must catch this exception and reset or fail the phase cleanly.

Now Semaphore, which bounds concurrency — the canonical connection-pool or rate-limiter:

Semaphore permits = new Semaphore(10, /*fair*/ true);
void call() throws InterruptedException {
    permits.acquire();
    try { doRemoteCall(); } finally { permits.release(); } // release in finally, always
}
The number-one semaphore bug: over-releasing

release() is not validated against a prior acquire(). An extra release() in some code path silently raises the permit count — the semaphore climbs from 10 permits to 11 and beyond, and the entire concurrency limit is destroyed, without you seeing any error. Rule: acquire outside the try, release in finally, exactly once.

And finally Phaser: it generalizes both latch and barrier. Parties can register/deregister dynamically, and it supports multiple phases without reconstruction — ideal for fork/join-style staged pipelines where the number of participants changes during execution.


The three classic problems

Every concurrency textbook has three famous puzzles, each illuminating one of the Coffman conditions. We already solved the first.

Bounded buffer — solved above (producer-consumer).

Readers–writers

A public bulletin board

Picture a bulletin board: a thousand people can read it at the same time with no trouble, but when one person wants to write or erase something, everyone else must step back so they're alone. Reading is shared, writing is exclusive.

Many readers may share; a writer needs exclusivity. ReentrantReadWriteLock handles it, but naive use starves writers under constant read traffic (if there's always some reader reading, the writer never gets its turn).

ReentrantReadWriteLock rw = new ReentrantReadWriteLock(true); // fair → no writer starvation
Lock r = rw.readLock(), w = rw.writeLock();

Object read()  { r.lock(); try { return data; } finally { r.unlock(); } }
void  write(Object x) { w.lock(); try { data = x; } finally { w.unlock(); } }
You may downgrade, but you may not upgrade

You may downgrade: hold the write lock, acquire the read lock, then release the write lock — this is safe. But you may not upgrade: hold the read lock and try to acquire the write lock — because the writer must wait for all readers to leave, including you, who still hasn't released your read lock. The result is self-deadlock.

For read-dominated data, StampedLock has a superb trick: optimistic read, which takes no lock at all on the happy path.

StampedLock sl = new StampedLock();
double distanceFromOrigin() {
    long stamp = sl.tryOptimisticRead();      // no lock taken
    double cx = x, cy = y;                     // read fields
    if (!sl.validate(stamp)) {                 // a writer intervened?
        stamp = sl.readLock();                 // fall back to a real read lock
        try { cx = x; cy = y; } finally { sl.unlockRead(stamp); }
    }
    return Math.sqrt(cx * cx + cy * cy);
}

The logic is: grab a "stamp," read without locking, then ask "did a writer come by in the meantime?" If not, your read was valid and it cost nothing. If so, fall back to a real read lock. Note: StampedLock is not reentrant and not Condition-capable — respect those limits.

Dining philosophers

The philosophers' dinner table

Five philosophers sit around a round table with one fork between each pair — five forks total. Each philosopher needs both neighboring forks to eat. If they all pick up their left fork at once, everyone holds one fork and waits for the right fork held by a neighbor — and everyone stays hungry. It's a live demonstration of circular wait.

The naive "grab left, then right" deadlocks when all grab left simultaneously. We have two clean fixes, each breaking a different Coffman condition:

// Fix A: break symmetry — one philosopher picks up right-first (resource ordering)
void dine(int id, Lock left, Lock right) {
    Lock first = (id == LAST) ? right : left;   // one philosopher inverts
    Lock second = (id == LAST) ? left : right;
    first.lock();
    try { second.lock();
        try { eat(); } finally { second.unlock(); }
    } finally { first.unlock(); }
}
// Fix B: limit concurrency to N-1 seated philosophers via a semaphore
Semaphore seats = new Semaphore(PHILOSOPHERS - 1);
void dine(...) throws InterruptedException {
    seats.acquire();               // at most 4 of 5 may contend for forks
    try { left.lock(); right.lock();
        try { eat(); } finally { right.unlock(); left.unlock(); }
    } finally { seats.release(); }
}
Two fixes, two broken conditions

Fix A breaks circular wait by inverting one philosopher's order (global ordering / asymmetry). Fix B breaks hold-and-wait by allowing at most N-1 seated — because with 4 people for 5 forks, at least one person can always grab both forks, and a full deadlock becomes arithmetically impossible. This shows how breaking any one of the four Coffman conditions is enough.


Immutability and confinement as strategies

We've learned plenty of tools for managing sharing. But the best strategy is to have no shared mutable state at all.

The cheapest concurrency is no shared mutable state

If data is shared but immutable, or not shared at all, no lock is needed and no race is possible. Two structural strategies — immutability and confinement — eliminate whole bug classes at the root.

Immutability.

A carved stone tablet

An immutable object is like a carved stone tablet: once made, nobody can change it. A thousand people can read it at once with no danger, because nobody writes. If you want something new, you carve a fresh tablet.

An object whose fields are all final and never escape during construction is safely published through the JMM's final-field guarantee and can be shared freely without synchronization. Java records make this natural:

record Money(long cents, String currency) {          // deeply immutable
    Money add(Money o) {                              // returns a NEW instance
        if (!currency.equals(o.currency)) throw new IllegalArgumentException();
        return new Money(cents + o.cents, currency);
    }
}
`final` protects the reference, not the object behind it

A record holding a List is not immutable unless you defensively copy it into an unmodifiable list — otherwise someone can mutate the list's contents even though the reference itself is final. And a final field is only safely published if this didn't escape during the constructor (e.g., if you didn't register yourself as a listener before construction finished).

Confinement. Keep data on one thread so no synchronization is needed at all. It has three forms:

  • Thread confinement via ThreadLocal — each thread gets its own copy. But beware leaks: on a thread pool, a ThreadLocal you don't remove() lives as long as the worker thread and can pin large objects or classloaders — a classic web-app memory leak. Always remove() in a finally.
  • Stack confinement — local variables and objects that never escape a method are automatically thread-safe, because each thread has its own stack. Prefer this by default; it's free.
  • Instance confinement — guard mutable state behind an object's own lock and never let a reference escape. This is the Java monitor pattern done deliberately and consciously.
Virtual threads (Java 21) don't change the rules

Virtual threads (Java 21, Thread.ofVirtual()) make blocking cheap so you can write straightforward blocking, confined-per-request code at massive scale. But a shared mutable object is just as unsafe from a virtual thread as from a platform thread, and the JMM rules are unchanged. Note: pinning occurs if a virtual thread blocks inside a synchronized block; on Java 21 prefer ReentrantLock for that reason (this limitation is largely resolved in Java 24+). And never pool virtual threads.


Concurrency testing tips

You cannot test correctness *into* a concurrent design

Concurrency bugs are non-deterministic; a single unit test might go green a thousand times and blow up on run one-thousand-and-one under real load. Ordinary tests give false confidence.

Techniques that actually find these bugs:

  • jcstress — the OpenJDK harness built specifically to expose JMM/reordering bugs by running billions of interleavings and classifying outcomes. Use it for any low-level lock-free code you write.
  • Stress with contention: run many threads (> cores) in a tight loop for seconds, use a CyclicBarrier to make them all start at the exact same instant (maximizing overlap), and assert an invariant afterward.
  • -Xint / -XX:-TieredCompilation and running on ARM/weak-memory hardware surface reordering that x86's strong memory model hides.
  • Deadlock watchdog in tests: a background ThreadMXBean.findDeadlockedThreads() poll that fails the test the moment a cycle appears, instead of hanging CI.
  • Thread.sleep in tests is a smell — it makes tests slow and still flaky. Use latches/barriers to express the actual ordering you want to force.
  • Fuzz the scheduler: tools like Thread.yield() injection, or JPF (Java PathFinder) model checking for exhaustive small-scope interleavings.

The honest truth: you reason correctness in with happens-before, keep the shared surface tiny, and use tests only to catch regressions.


Interview Questions

Now it's time to gather everything into real interview questions. Read each with its full answer and try to answer it yourself before you look.

1) State the four Coffman conditions and give one prevention for each.

Mutual exclusion (use immutable/lock-free), hold-and-wait (acquire all at once), no preemption (tryLock + timeout), circular wait (global lock ordering). Breaking any single one prevents deadlock; circular-wait removal via lock ordering is the most practical in most systems.

2) Difference between deadlock, livelock, and starvation.

Deadlock: threads blocked forever in a wait cycle. Livelock: threads actively running and reacting but making no progress (symmetric retry). Starvation: a thread makes no progress because others keep winning the resource. Deadlock/livelock are typically symmetry/ordering problems; starvation is a fairness problem.

3) (Gotcha) Why must condition waits use `while`, not `if`?

Spurious wakeups are permitted by the spec, and even without them, another thread may consume the condition between your wakeup and your reacquiring the lock. Re-checking the predicate in a loop is the only correct pattern. An if produces intermittent corruption that's nearly impossible to reproduce.

4) `notify` vs `notifyAll` — when is `notify` safe?

notify wakes one arbitrary waiter. It's safe only when all waiters are interchangeable (wait on the same condition and any one making progress is fine) and you signal exactly one available unit. If waiters wait on different predicates on the same monitor, notify can wake the wrong one and cause a lost-wakeup hang; use notifyAll or, better, distinct Condition objects.

5) (Gotcha) What does this print?
List<Integer> list = new ArrayList<>();
IntStream.range(0, 4).parallel().forEach(list::add);
System.out.println(list.size());

Undefined — anything from 1 to 4, or an ArrayIndexOutOfBoundsException/NullPointerException. ArrayList is not thread-safe; concurrent add races on size and the backing array. Fix: Collections.synchronizedList, a concurrent collection, or .collect(Collectors.toList()) on the stream.

6) (Gotcha — find the bug)
if (!map.containsKey(k)) map.put(k, compute(k)); // map is ConcurrentHashMap

Check-then-act race: two threads both see the key absent and both compute/put. Individual ConcurrentHashMap ops are atomic, but the compound op is not. Fix: map.computeIfAbsent(k, this::compute).

7) When would you choose `LongAdder` over `AtomicLong`?

High write contention with infrequent reads. AtomicLong's single CAS location becomes a hotspot; LongAdder stripes across cells and sums on read, trading exact-at-all-times reads and memory for far higher write throughput. For low contention or when you read the value constantly, AtomicLong is simpler and fine.

8) Why prefer a bounded BlockingQueue over an unbounded one?

Backpressure. A bounded queue makes producers block when full, propagating slowness upstream. An unbounded queue absorbs a load spike into unbounded heap growth and eventually OutOfMemoryError, turning a latency problem into an outage. Capacity is a design parameter, not a nuisance.

9) (Hard) Explain the JMM `final`-field guarantee and how a data race can still break it.

If an object's final fields are set in the constructor and this does not escape during construction, any thread that sees a reference to the object is guaranteed to see the correctly initialized final fields, without synchronization. It breaks if this escapes the constructor (e.g., registering a listener before construction finishes) — then another thread can observe partially built state.

10) Can two threads deadlock with a single lock?

Not on a reentrant lock re-acquired by the same thread. But yes across two threads if one holds lock A and calls a method that needs lock B while the other holds B and needs A — the "single object" can still be part of a two-lock cycle. Also, a non-reentrant lock (like StampedLock) can self-deadlock if the same thread re-locks it.

11) (Gotcha) A `CyclicBarrier(3)` has two threads waiting and a third times out on `await`. What happens to the two waiters?

They receive a BrokenBarrierException. A barrier is all-or-nothing: a timeout, interrupt, or failed action breaks it for everyone currently waiting. Robust code catches this and resets or fails the phase cleanly.

12) Semaphore correctness pitfall?

release() is not validated against prior acquire(). An extra release — often on an exception path where you release() in a finally that also ran without a matching acquire() — silently raises the permit count and destroys the concurrency limit. Acquire outside the try, release in finally, exactly once.

13) (Hard) Why can `ThreadLocal` leak on a thread pool, and how do you prevent it?

Pool threads are long-lived, so a value you set persists across unrelated tasks and pins whatever it references (large buffers, classloaders in app servers) until the thread dies. Prevent by wrapping usage in try/finally { threadLocal.remove(); } at the task boundary. InheritableThreadLocal compounds this across spawned threads.

14) How do virtual threads (Java 21) change concurrency design, and what stays the same?

They make blocking cheap, so you can use simple synchronous, thread-per-request, confined code at scale instead of reactive callback chains — fewer thread pools to tune. What stays the same: shared mutable state is exactly as dangerous, and the JMM rules are unchanged. Watch for pinning when blocking inside synchronized (prefer ReentrantLock on 21) and never pool virtual threads.

15) (Hard) Give a scenario where adding synchronization *creates* a deadlock, and how you'd detect it.

You wrap two previously-independent methods in synchronized to fix a race; now a call graph acquires lock A→B on one path and B→A on another, forming a cycle. Detect with a thread dump (jstack prints the deadlock and cycle) or ThreadMXBean.findDeadlockedThreads() in a watchdog. Fix by imposing a global lock order or shrinking the critical section so nested locking disappears.


Senior notes & advanced edge cases

So far we built the map of concurrency bugs and their defensive patterns. But what separates a real senior from a merely good programmer is not knowing synchronized and BlockingQueue — everyone knows those. The difference lives where code breaks in production under real load: the thread pool that silently ignores your extra threads, the swallowed InterruptedException that quietly kills cancellation for the whole system, the CompletableFuture blocking on the wrong pool, and the bugs that live at the level of the CPU cache line. This section is exactly that layer.

Roadmap for this section

First we go to the beating heart of every Java service: ThreadPoolExecutor and its unbounded-queue trap. Then the interruption protocol (why "swallowing InterruptedException" is a crime). Then CompletableFuture, the commonPool trap, and the Memoizer pattern that defeats cache stampede. Then double-checked locking and the holder idiom. Then down to the hardware: false sharing. Then the ABA problem in CAS. And finally the modern-Java concurrency map (2025–2026): the death of biased locking, the JDK 24 pinning fix, and Scoped Values. Then a set of hard senior questions.

The thread pool — where most production incidents are born

Most of the concurrency code you actually write never touches a raw Thread. You build an ExecutorService and hand it tasks. Behind Executors.newFixedThreadPool(...) sits a ThreadPoolExecutor, and that class has a task-admission logic that will ambush you one day if you don't know it.

A restaurant with cooks, waiting chairs, and on-call cooks

corePoolSize is the always-on cooks. The queue is the chairs where orders wait. maximumPoolSize is the ceiling of extra cooks you can call in during a rush. The golden rule, which is where everyone gets surprised: an extra cook is hired only once the queue is full — not sooner.

The admission order for a new task is precisely this, and its order is exactly where people go wrong:

1) if active threads < corePoolSize  → create a core thread and run it
2) else, offer the task to the queue (queue.offer)
3) if the queue is full  → create extra threads up to maximumPoolSize
4) if that too is full   → invoke the RejectedExecutionHandler

The diagram below shows this decision flow (Task admission flow):

flowchart TD
  A[New task submitted] --> B{active < corePoolSize?}
  B -- yes --> C[Start core thread]
  B -- no --> D{queue.offer succeeds?}
  D -- yes --> E[Task waits in queue]
  D -- no --> F{active < maximumPoolSize?}
  F -- yes --> G[Start extra thread]
  F -- no --> H[RejectedExecutionHandler]

Now the killer trap. Step 2 says "while the queue has room, enqueue." If your queue is unbounded (like a capacity-less LinkedBlockingQueue — which is exactly what newFixedThreadPool builds), offer never fails. That means step 3 never runs, and your maximumPoolSize is a purely decorative number.

`maximumPoolSize` with an unbounded queue is a complete lie

If you build a ThreadPoolExecutor with corePoolSize=10, maximumPoolSize=100 and a capacity-less LinkedBlockingQueue, your system will never exceed 10 threads, no matter how much load arrives. Tasks pile up silently in the queue, latency explodes, and eventually the heap fills and you get an OutOfMemoryError. This is precisely why teams stop using Executors.* and instead new ThreadPoolExecutor(...) directly with a bounded queue and a saturation policy.

The saturation policies (invoked when both queue and threads are full) must be chosen deliberately:

  • AbortPolicy (default): throws RejectedExecutionException — the caller finds out.
  • CallerRunsPolicy: runs the task in the calling thread itself. This is a brilliant automatic backpressure brake — the web thread that wanted to submit is forced to do the work itself and cannot accept new tasks while it does.
  • DiscardPolicy / DiscardOldestPolicy: silently drops work — almost always wrong, because you lose data with no signal.
Pool sizing formula (Brian Goetz)

For compute-bound work: threads ≈ number of cores + 1. For I/O-bound work, the classic formula is: N = N_cpu × U × (1 + W/C) where U is target utilization (0 to 1), W is wait time (I/O), and C is compute time per task. Senior insight: the larger the W/C ratio (more I/O), the more threads you need. But with the arrival of virtual threads, this arithmetic is largely obsolete for I/O-bound work — you use Executors.newVirtualThreadPerTaskExecutor() and stop tuning pool size altogether.

And two shutdown facts I constantly see wrong in review:

`shutdown()` does not wait, and `Future.get` wraps the exception

executor.shutdown() only flags "stop accepting new tasks" and returns immediately; to actually wait you must then call awaitTermination(...). shutdownNow() interrupts the threads, but only works if your code respects interrupts (next section). And when a task throws, future.get() wraps it in an ExecutionException — you must unwrap e.getCause(). Worse: if you submit via execute(...) (not submit) and the task throws, the exception goes silently to the UncaughtExceptionHandler, and you'll see nothing in your logs unless you've set one.

Interruption is a protocol, not a kill switch

The biggest misconception among young seniors is thinking thread.interrupt() "kills" a thread. It doesn't. Interrupt merely sets a boolean flag on the thread; it is a polite cancellation request that the code itself must respond to. The entire cancellation mechanism in Java is built on this contract.

Crime number one: swallowing `InterruptedException`

You see this everywhere:

try {
    Thread.sleep(1000);
} catch (InterruptedException e) {
    // ... nothing
}

When a blocking method throws InterruptedException, the JVM clears the interrupt flag. If you catch the exception and do nothing, the cancellation signal has vanished forever — higher layers no longer know this thread should die, and a shutdownNow() or a timeout becomes a no-op. The result: threads that never finish, piled up in your thread dump.

The correct rule has two cases. If you can propagate the InterruptedException upward, do so (let your method throw it). If you cannot (e.g. inside a Runnable whose signature forbids it), you must restore the flag:

try {
    queue.take();
} catch (InterruptedException e) {
    Thread.currentThread().interrupt(); // restore the flag
    return; // and exit the loop/task
}

And for long compute loops that contain no blocking method (so never receive an InterruptedException), you must poll the flag yourself:

while (!Thread.currentThread().isInterrupted()) {
    doOneChunkOfWork();
}
Interruption is cooperative

Nobody forcibly stops your thread (Thread.stop() has been deprecated for years and is dangerous because it drops locks mid-operation). Cancellation only works when all the code on the path — yours and your libraries' — respects the interrupt flag. One bad library that swallows the exception breaks cancellation for the whole system.

CompletableFuture: async composition and its traps

The main chapter covered Future and executors but not the modern async-composition tool — CompletableFuture (since Java 8). This is what lets you build a non-blocking chain of stages instead of blocking on future.get().

The first distinction interviewers ask: thenApply vs thenCompose.

// thenApply: synchronous function, value → value
CompletableFuture<Integer> len = fetchUser(id).thenApply(User::name).thenApply(String::length);

// thenCompose: a function that itself returns a CompletableFuture → flattens it (flatMap)
CompletableFuture<Order> order = fetchUser(id).thenCompose(u -> fetchLatestOrder(u)); // no nested future

If you pass thenApply a function that itself returns a CompletableFuture, you get a CompletableFuture<CompletableFuture<T>> — exactly like map vs flatMap on a stream. thenCompose is the flatMap.

The silent trap: `commonPool` and the `*Async` methods

Every method without the Async suffix runs on the same thread that completed the previous stage. The *Async methods with no Executor argument run on ForkJoinPool.commonPool(). The problem: commonPool defaults to (cores − 1) threads and is shared across the whole JVM — the same pool parallelStream() uses. If you block on I/O inside a stage, you starve the commonPool threads, and suddenly unrelated parallelStream() calls elsewhere in the app slow down. Rule: never run blocking or I/O work on the commonPool; always pass an explicit Executor to *Async.

Exception handling has its own trap: in a chain, an exception skips the following stages until it reaches exceptionally/handle.

fetchUser(id)
    .thenApply(this::risky)
    .exceptionally(ex -> User.GUEST)      // catches only the error path, supplies a value
    .thenAccept(this::render);
// handle(value, ex) catches both paths; whenComplete is a side-effect and does not alter the exception

Since Java 9, orTimeout(...) and completeOnTimeout(...) were added so you no longer have to hand-roll timeouts.

The Memoizer pattern: defeating cache stampede

The chapter used computeIfAbsent to fix the get-then-put race. But there's a deeper problem: if the computation is expensive and slow and 100 threads want the same absent key at once, does it compute 100 times? Goetz's classic solution is to cache the CompletableFuture itself in the map, not the value:

ConcurrentHashMap<K, CompletableFuture<V>> cache = new ConcurrentHashMap<>();
V get(K key) {
    CompletableFuture<V> f = cache.computeIfAbsent(key, k -> CompletableFuture.supplyAsync(() -> compute(k)));
    return f.join();
}

Now the first thread creates the future and everyone else gets the same in-flight future and waits on it — the computation runs exactly once. This kills the "thundering herd / cache stampede." (Note: if the computation fails, remove the failed cached future so a retry becomes possible.)

Double-checked locking and the holder idiom

A classic pattern the chapter didn't cover: lazy, thread-safe initialization without taking a lock on the hot path. The naive version was broken for years:

private Helper helper; // BUG: no volatile
Helper get() {
    if (helper == null) {                 // check 1 (no lock)
        synchronized (this) {
            if (helper == null)           // check 2 (with lock)
                helper = new Helper();
        }
    }
    return helper;
}
Why `volatile` is critical here

helper = new Helper() is not one atomic operation: it's memory allocation, running the constructor, and assigning the reference. Without volatile, the JMM permits these to be reordered — meaning the reference can be published before the constructor finishes. A second thread that passes the lock-free first check gets a non-null reference to a half-constructed object. The fix: private volatile Helper helper; — which creates a happens-before edge and forbids the reordering.

But there's a cleaner, better way that needs no volatile and no double-check at all — the initialization-on-demand holder idiom:

class Config {
    private Config() { /* expensive */ }
    private static class Holder { static final Config INSTANCE = new Config(); }
    static Config get() { return Holder.INSTANCE; } // lazy + thread-safe, free
}
Why the holder beats any lock

The JVM guarantees a class is initialized exactly once, thread-safely (under an internal class-init lock). The Holder class isn't loaded until the first reference to Holder.INSTANCE — so initialization is lazy — and mutual exclusion is handed to you for free by the classloader. No volatile, no synchronized, zero cost on the hot path. For singletons this is ideal (or a single-element enum).

The hardware layer: false sharing

Some performance bugs have nothing to do with your code's logic and live in the CPU's cache line. The processor moves memory not byte-by-byte but in 64-byte blocks (cache lines).

Two people and a shared whiteboard

Two workers write on two different corners of one whiteboard. Logically they don't interfere. But the whiteboard is so small that every time one writes, the system has to "invalidate" the whole board for the other and re-copy it. They share no data but they share a physical location — and that alone slows them down.

If two independent variables that different threads write to happen to land on the same cache line, every write by one invalidates the other's cache, and the CPU's coherency protocol keeps ping-ponging that line between cores. The code is correct but perhaps 5–10× slower. It's called false sharing because there's no actual sharing.

How you spot it and cure it in practice

If you have a hot counter and you're suspicious, Java has @jdk.internal.vm.annotation.Contended (and its public variants) which pads a field away from the rest — you must run the JVM with -XX:-RestrictContended. But the healthier route is usually to use tools that pad themselves: this is exactly why LongAdder is faster than an array of AtomicLong — its cells are padded. This is a great interview question for low-latency roles (fintech, trading).

CAS and the ABA problem

The chapter covered AtomicLong and CAS but not one of the subtlest bugs in non-blocking algorithms: the ABA problem. CAS says "if the value is still A, change it to B." But what if, between your read and your CAS, the value went from A to X and back to A? Your CAS succeeds, because it only sees the value, not the history.

A door lock that was swapped and reinstalled

You leave the house and check the door with your key: lock "A". You return, it's still lock "A", so you conclude "nothing has changed." But in between, someone removed the whole lock, emptied the house, and reinstalled an apparently identical lock. The value is the same, but the world under your feet has changed.

In lock-free linked structures (like a Treiber stack that reuses nodes from a free-list), ABA can reattach a freed node and corrupt the structure. The fix: attach a version number (stamp) to each value.

AtomicStampedReference<Node> top = new AtomicStampedReference<>(head, 0);
int[] stampHolder = new int[1];
Node cur = top.get(stampHolder);
// ... CAS with the value *and* the next stamp; even if the value returns to cur, the stamp differs
top.compareAndSet(cur, next, stampHolder[0], stampHolder[0] + 1);

AtomicStampedReference makes the value and an int counter atomic together; now A→X→A no longer fools you because the stamp has advanced. (AtomicMarkableReference is the boolean variant for logical-deletion marking.)

The modern-Java concurrency map (2025–2026)

A few big changes every senior should know in a 2026 interview:

Biased locking is dead

For a long time the JVM had an optimization called biased locking (assuming a lock is usually always held by the same thread). It was disabled by default in JDK 15 via JEP 374 and later removed entirely, because with today's heavily concurrent code it did more harm than good. Practical consequence: an uncontended synchronized today is slightly more expensive than it used to be — one more reason to keep critical sections small.

Pinning was fixed in JDK 24 (JEP 491)

The chapter correctly said that in Java 21, if a virtual thread blocks inside synchronized, it gets pinned to its carrier thread and breaks scalability. Big news: JEP 491 in JDK 24 fixed this almost entirely — the monitor is now associated with the virtual thread itself, so a thread can block even inside synchronized and still release its carrier. That means the advice "prefer ReentrantLock over synchronized for virtual threads" is mostly a pre-JDK-24 concern. (Pinning still remains in native frames and class initializers, but those are rare.)

Scoped Values replaced ThreadLocal (JDK 25, JEP 506)

The chapter rightly warned about ThreadLocal leaks on pools. Modern Java has a better replacement: Scoped Values, finalized in JDK 25. You share an immutable value for the duration of an operation and all its subtasks (and child threads), and it's cleared automatically at the end of the scope — no manual remove(), no leak, and cheaper than ThreadLocal, especially with millions of virtual threads:

private static final ScopedValue<User> CURRENT = ScopedValue.newInstance();
ScopedValue.where(CURRENT, user).run(() -> handleRequest()); // readable in scope, not outside

Alongside it, Structured Concurrency (StructuredTaskScope) — still preview (its fifth preview in JDK 25) — lets you manage a group of subtasks as one unit: either all succeed, or all are cancelled together — the end of leaked, orphaned threads.

Hard senior interview questions

1) You built a `ThreadPoolExecutor` with core=5, max=50 and an unbounded `LinkedBlockingQueue`. Under heavy load, how many threads will you have?

Exactly 5. With an unbounded queue, offer never fails, so the pool logic never reaches the stage of creating extra threads (up to max); maximumPoolSize is completely inert. Tasks silently pile up in the queue until OOM. Correct: use a bounded queue (ArrayBlockingQueue) together with a RejectedExecutionHandler like CallerRunsPolicy, so you actually reach max and get real backpressure.

2) (Trap) What's wrong with this code?
try { doBlockingWork(); }
catch (InterruptedException e) { log.warn("interrupted"); }

It swallows the cancellation signal. When InterruptedException was thrown, the JVM cleared the interrupt flag; this code neither restores it nor propagates it upward. As a result, higher layers (and shutdownNow()/timeouts) no longer know the thread should stop, and the thread lives forever. Correct: either throw the exception, or call Thread.currentThread().interrupt() and exit the task.

3) What is the difference between `thenApply` and `thenCompose` in `CompletableFuture`, and why does it matter?

thenApply is like map: it takes a synchronous T → U function. thenCompose is like flatMap: it takes a T → CompletableFuture<U> function and flattens the result. If you pass thenApply a function that itself returns a future, you get a nested CompletableFuture<CompletableFuture<U>> that's a nightmare to work with. Wherever the next stage is itself asynchronous (another service call), thenCompose is the right choice.

4) (Hard) Why is running blocking work on `CompletableFuture.supplyAsync(...)` without an Executor dangerous?

Because it defaults to ForkJoinPool.commonPool(), which (a) has only cores − 1 threads and (b) is shared across the whole JVM — the same pool parallelStream() uses. If you block on I/O inside it, you occupy the commonPool's limited threads and starve unrelated parallel work elsewhere in the app; worst case, with recursive fork/join tasks, a self-deadlock. Correct: always pass an explicit dedicated Executor (or a virtual-thread executor) to the *Async variant.

5) (Trap) How does double-checked locking break without `volatile`?

instance = new Helper() is three steps: allocate, run the constructor, assign. Without volatile, the JMM permits reordering, so the reference may be published before the constructor finishes. A second thread that sees the lock-free first check gets a non-null reference to a half-constructed object and uses it. The field must be volatile to create a happens-before edge. Better still: the initialization-on-demand holder idiom, which sidesteps the trap entirely.

6) (Hard) What is false sharing and why does it slow down perfectly correct code?

The CPU moves memory in 64-byte cache lines. If two independent variables that different threads write to land on the same cache line, each write by one invalidates the other's cached copy, and the cache-coherency protocol keeps ping-ponging that line between cores. Logically there's no sharing at all, but performance drops several-fold. Cure: padding (like @Contended) or using padded structures like LongAdder. A great example of "correctness ≠ performance."

7) Explain the ABA problem in CAS and how you fix it.

CAS only sees the value, not the history. If a value goes A → B → back to A, your CAS succeeds as if nothing changed — even though an invariant you relied on may have been violated (e.g. a freed and reused node in a lock-free stack). Fix: attach a version number to each value with AtomicStampedReference, which compares the value and stamp together atomically; since the stamp always advances, an A→X→A return no longer fools you.

8) In 2025/2026 Java, is the advice "use `ReentrantLock` instead of `synchronized` for virtual threads" still correct?

It's mostly a pre-JDK-24 concern. In Java 21, a virtual thread blocking inside synchronized got pinned to its carrier thread and broke scalability. But JEP 491 in JDK 24 fixed this: the monitor is now associated with the virtual thread itself, so the thread can unmount even inside synchronized. So on JDK 24+, synchronized no longer pins (except rare native-frame and class-initializer cases). This shows that "best practices" are versioned and you must know which JDK you're speaking about.

Senior notes in a nutshell

Behind every service is a ThreadPoolExecutor: an unbounded queue turns maximumPoolSize into a decorative lie — use a bounded queue + CallerRunsPolicy. Interruption is a cooperative protocol, not a kill switch; never swallow InterruptedException — either throw it or restore the flag. In CompletableFuture, thenCompose is the flatMap, and never block on the commonPool; the Memoizer pattern kills cache stampede by caching the future itself. For lazy init, prefer the holder idiom over double-checked locking (which breaks without volatile). Some performance drops are cache-line false sharing, not logic. CAS is vulnerable to ABA — version it with AtomicStampedReference. And know the modern map: biased locking is dead, pinning was fixed in JDK 24, and Scoped Values are the safe successor to ThreadLocal.

In a nutshell

Every concurrency bug either violates safety (wrong answer: races, check-then-act) or liveness (a hang: deadlock, livelock, starvation) — and the fixes pull in opposite directions, so stay balanced on that narrow ridge. Underneath everything is the JMM's happens-before relationship; without it you have no guarantees. Deadlock needs all four Coffman conditions — break just one, usually with global lock ordering or tryLock+timeout (and that critical finally). Cure livelock with asymmetry (randomized backoff), starvation with fair locks. For races, use a single atomic operation: computeIfAbsent, AtomicLong/LongAdder. For producer-consumer, use a bounded BlockingQueue (backpressure) and a poison pill; in the hand-rolled version always while not if, with two Conditions. Know the coordination tools (one-shot latch, reusable-and-breakable barrier, semaphore with its over-release danger, dynamic phaser). And above all: the cheapest concurrency is immutability and confinement — if nothing is shared and mutable, there's no bug to have. You reason correctness in; tests only catch regressions.