Concurrency · همزمانی متوسطIntermediate ~68 دقیقه مطالعه~59 min read

نخ‌ها، Runnable/Callable و ExecutorهاThreads, Runnable/Callable & Executors

از صفر تا سنیور یاد می‌گیری که نخ چیست، چطور با Runnable/Callable/Future کار بسپاری، چرا Executorها را دوست داری، ThreadPoolExecutor را چطور بسازی و اندازه بگیری، چطور باوقار خاموش کنی، لغو تعاونی و ThreadLocal را رام کنی، و کجا نخ‌های مجازی همه‌چیز را عوض می‌کنند.A zero-to-senior walkthrough of what a thread really is, how to hand off work with Runnable/Callable/Future, why Executors exist, how to build and size a ThreadPoolExecutor, shut it down gracefully, tame cooperative cancellation and ThreadLocal, and where virtual threads rewrite the rules.


خب، بنشین کنار من. می‌خواهیم دربارهٔ یکی از آن موضوع‌هایی حرف بزنیم که همه در مصاحبه ازش می‌ترسند و در تولید (production) ازش خون‌دل می‌خورند: همزمانی (concurrency). اما قول می‌دهم اگر سه چیز را درست بفهمی، این موضوع از حالت «جادوی ترسناک» درمی‌آید و می‌شود چیزی که کف دستت است. تمام این فصل، همان سه چیز است که با آجرهای کوچک، یکی‌یکی، از پایه می‌سازیمشان.

قبل از هر چیز یک تصویر ذهنی بگیر: یک نخ (thread) مثل یک کارگر است که می‌تواند یک کار را دنبال کند. اگر یک کارگر داشته باشی، کارها پشت سر هم انجام می‌شوند. اگر چند کارگر داشته باشی، کارها می‌توانند هم‌زمان پیش بروند. تمام دعوای این فصل سر این است که این کارگرها را چطور بسازی، چطور بهشان کار بدهی، و چطور مؤدبانه بگویی «دیگر بس است، برو خانه».

نقشهٔ راه

در این فصل این مسیر را می‌رویم:

  1. نخ واقعاً چیست و چرا ساختنش گران است (و همین گرانی، دلیل وجود Executorهاست).
  2. چرخهٔ حیات نخ — شش حالتی که یک نخ می‌تواند در آن باشد.
  3. واحد کار: Runnable، Callable و دستگیرهٔ نتیجه یعنی Future.
  4. چارچوب Executor و قلبش ThreadPoolExecutor با هفت پارامترش.
  5. ریاضیِ اندازه‌گذاری استخر — همان جایی که سنیورها امتیاز می‌گیرند.
  6. زمان‌بندی (ScheduledExecutorService) و خاموش‌سازیِ باوقار.
  7. لغو تعاونی (InterruptedException) و ThreadLocal که در استخرها نشت می‌کند.
  8. نخ‌های daemon و در پایان نخ‌های مجازی (Java 21+) که کل ریاضیِ اندازه‌گذاری را دود می‌کنند. در انتها پانزده پرسش مصاحبه با پاسخ کامل داریم. آماده‌ای؟ برویم.

بخش ۰ — واژه‌هایی که باید از پیش بشناسی

قبل از اینکه بپریم وسط کد، بیا چند واژه را که مدام به کار می‌بریم، با یک تشبیه کوچک جا بیندازیم. اگر اینها ته ذهنت باشد، بقیهٔ فصل مثل آب خوردن می‌شود.

رستوران و آشپزها

یک رستوران را تصور کن. آشپز همان نخ است: یک نفر که می‌تواند یک سفارش را از اول تا آخر بپزد. مدیر سالن که تصمیم می‌گیرد کدام آشپز الان کار کند همان زمان‌بند سیستم‌عامل (OS scheduler) است — چون آشپزخانه فقط چند اجاق (CPU/هسته) دارد و همه نمی‌توانند هم‌زمان بپزند، مدیر نوبت می‌دهد. میز کار شخصی هر آشپز با تخته و چاقو و ادویه‌اش، همان پشته (stack) است: حافظه‌ای که فقط مال خودِ آن نخ است. رفتن تا انبار و برگشتن برای گرفتن یک آشپز جدید، همان فراخوان سیستمی (system call) است: کند و پرهزینه. برای همین رستوران‌های حرفه‌ای آشپز را وسط شیفت استخدام و اخراج نمی‌کنند؛ یک تیم ثابت دارند و سفارش‌ها را بینشان پخش می‌کنند. این دقیقاً همان چیزی است که Executor برایت انجام می‌دهد.

  • نخ (thread): واحد اجرای مستقل. یک نخ یک دنباله از دستورها را جلو می‌برد.
  • زمان‌بند سیستم‌عامل (OS scheduler): بخشی از سیستم‌عامل که تصمیم می‌گیرد کدام نخ روی کدام هسته و چه مدت اجرا شود. تو مستقیم کنترلش نمی‌کنی، فقط درخواست می‌کنی.
  • پشته (stack): حافظهٔ خصوصیِ هر نخ برای متغیرهای محلی و زنجیرهٔ فراخوان متدها. هر نخ پشتهٔ خودش را دارد (معمولاً حدود ۱ مگابایت رزرو).
  • فراخوان سیستمی (system call): وقتی برنامهٔ تو از هستهٔ سیستم‌عامل چیزی می‌خواهد (مثلاً «یک نخ تازه بساز»). این‌ها نسبتاً کندند.
  • مسدود شدن (blocking): وقتی یک نخ می‌ایستد و منتظر می‌ماند تا اتفاقی بیفتد — مثلاً منتظر پاسخ دیتابیس. در این مدت آن نخ کاری نمی‌کند جز انتظار.
  • قفل مانیتور (monitor lock): قفلی که با کلیدواژهٔ synchronized روی یک شیء گرفته می‌شود تا فقط یک نخ هم‌زمان واردِ ناحیهٔ حساس شود. مثل کلید دستشویی که فقط یکی دستش است.
  • CPU-محور در برابر I/O-محور: کارِ CPU-محور یعنی نخ واقعاً محاسبه می‌کند (فشرده‌سازی، هش). کارِ I/O-محور یعنی نخ بیشتر منتظر چیزی بیرون است (دیتابیس، شبکه، دیسک). این تمایز، ستون فقرات اندازه‌گذاری استخر است.

خب، حالا که ابزارها دستمان است، برویم سراغ خودِ نخ.

مدل ذهنی: نخ چیست و چرا گران است

یک java.lang.Thread در جاوا در واقع پوششی نازک روی یک نخ سیستم‌عامل است — به این نخِ سیستم‌عاملی می‌گوییم نخ سکویی (platform thread)، تا بعداً از نخ مجازی جدایش کنیم.

حالا چرا می‌گویم گران است؟ وقتی یک نخ می‌سازی، سه اتفاق پرهزینه می‌افتد:

  1. یک پشتهٔ بومی بزرگ تخصیص داده می‌شود (معمولاً حدود ۱ مگابایت رزرو).
  2. نخ نزد زمان‌بند سیستم‌عامل ثبت می‌شود.
  3. همهٔ این‌ها یک فراخوان سیستمی می‌طلبد.
چرا اصلاً Executor داریم؟

همین گرانیِ ساختِ نخ، تمام دلیل وجود چارچوب Executor است. به‌جای اینکه برای هر کار یک آشپز تازه استخدام کنی (کُند و گران)، یک تیم کوچکِ ثابت از آشپزها داری و کارها را بینشان پخش می‌کنی. به این می‌گویند جدا کردن ثبت وظیفه از ساخت نخ (decouple task submission from thread creation). کار را می‌سپاری؛ اینکه کدام نخِ آماده اجرایش کند، دیگر دغدغهٔ تو نیست.

سه پرسش کل این فصل را قاب می‌گیرند. اگر جواب این سه را بلد باشی، همزمانی دیگر رازآلود نیست:

  1. واحد کار چیست؟ Runnable (بدون نتیجه، بدون استثنای checked)، Callable<V> (مقدار برمی‌گرداند و می‌تواند استثنای checked پرتاب کند)، یا زیرکلاس خامِ Thread (که تقریباً هیچ‌وقت پاسخ درست نیست).
  2. چه کسی و روی کدام نخ اجرایش می‌کند؟ یک Thread خام، یا یک ExecutorService که استخری (pool) پشتش است.
  3. چطور متوقف می‌شود؟ به‌صورت تعاونی، از طریق وقفه (interruption) — چون جاوا هیچ Thread.stop() امنِ پیش‌دستانه‌ای (preemptive) ندارد.

بیا همین سه را یکی‌یکی باز کنیم.

چرخهٔ حیات و حالت‌های نخ

هر نخ در طول عمرش بین چند «حالت» جابه‌جا می‌شود. جاوا این حالت‌ها را در enumی به نام Thread.State نگه می‌دارد. این enum حقیقتِ پایه است — حفظش کن، چون مصاحبه‌گر تقریباً همیشه تفاوتِ WAITING و BLOCKED را می‌کاود.

یک کارمند در طول روز

یک کارمند را تصور کن. NEW: تازه استخدام شده ولی هنوز سرِ کار نیامده. RUNNABLE: پشت میزش نشسته و کار می‌کند یا آمادهٔ کار است. BLOCKED: پشت درِ اتاق کنفرانس منتظر است تا نفر قبلی بیرون بیاید و کلید (قفل) را بدهد. WAITING: منتظر است تا همکارش بهش خبر بدهد، بدون هیچ مهلتی. TIMED_WAITING: همان انتظار، ولی با ساعت زنگ‌دار — «تا ساعت ۳ صبر می‌کنم». TERMINATED: کارش تمام شده و رفته خانه.

حالت معنا چطور به آن می‌رسی
NEW ساخته‌شده، هنوز start نشده new Thread(...) پیش از start()
RUNNABLE آمادهٔ اجرا یا در حال اجرا (JVM «آماده» و «روی CPU» را تفکیک نمی‌کند) بعد از start()
BLOCKED منتظر گرفتنِ قفل مانیتور (synchronized) رقابت بر سرِ یک مانیتور
WAITING انتظار نامحدود برای سیگنالِ نخ دیگر wait(), join(), LockSupport.park()
TIMED_WAITING همان، ولی با مهلت sleep(ms), wait(ms), join(ms), park(ns)
TERMINATED run() بازگشت یا پرتاب کرد وظیفه پایان یافت

حالا دو تلّه که دقیقاً همان جایی است که مصاحبه‌گر انگشت می‌گذارد:

نخِ منتظرِ I/O، RUNNABLE است نه BLOCKED

این یکی خیلی‌ها را غافلگیر می‌کند. نخی که روی یک عملیات I/O مسدود است (مثلاً وسط خواندن از یک سوکت شبکه منتظر داده مانده) در thread dump به‌صورت RUNNABLE نشان داده می‌شود، نه BLOCKED. چرا؟ چون این انسداد در سطح سیستم‌عامل رخ می‌دهد و JVM اصلاً آن را نمی‌بیند — از دید JVM آن نخ «قادر به اجرا» است. برچسب BLOCKED را جاوا فقط و فقط برای رقابت بر سرِ قفلِ مانیتورِ synchronized استفاده می‌کند. پس دفعهٔ بعد که thread dump دیدی و همه‌جا RUNNABLE بود، فریب نخور؛ ممکن است نخ‌ها در واقع منتظر شبکه باشند.

تلّهٔ کلاسیک: run() به‌جای start()

start() یک‌بارمصرف است. اگر start() را دوبار صدا بزنی، استثنای IllegalThreadStateException می‌گیری. و مهم‌تر: اگر مستقیم run() را صدا بزنی، اصلاً نخِ جدیدی راه نمی‌افتد — بدنه روی نخِ فعلی و به‌صورت هم‌ریسمانی (synchronous، یعنی همان لحظه و همان‌جا) اجرا می‌شود. یعنی هیچ همزمانی‌ای اتفاق نمی‌افتد. این یکی از رایج‌ترین باگ‌های تصادفیِ تازه‌کارهاست.

Runnable task = () -> System.out.println("on " + Thread.currentThread());
Thread t = new Thread(task);
t.run();    // باگ: روی نخِ *فعلی* اجرا می‌شود، بدون همزمانی
t.start();  // درست: نخ جدید می‌سازد

فرق start() و run() را این‌طور در ذهنت نگه دار: run() فقط یک متد عادی است که تو صدایش می‌زنی؛ start() است که به سیستم‌عامل می‌گوید «یک نخ تازه بساز و آن نخ، run() را صدا بزند».

Thread در برابر Runnable در برابر Callable/Future

حالا برویم سراغ پرسش اول: واحد کار چیست؟

یک قاعدهٔ طلایی: ترکیب (composition) را بر وراثت (inheritance) ترجیح بده. یعنی به‌جای اینکه کلاست را از Thread ارث‌بری کنی (extends Thread)، کارت را به‌شکل یک Runnable یا Callable بنویس و آن را بده به نخ. چرا؟ چون وقتی Thread را زیرکلاس می‌کنی، کارت را برای همیشه به یک نخ گره می‌زنی و دیگر نمی‌توانی همان کار را در یک استخر بازاستفاده کنی. اما یک Runnable یک بستهٔ کار مستقل است که هر نخی می‌تواند بازش کند و انجامش دهد.

برگهٔ سفارش در برابر آشپزِ اختصاصی

Runnable/Callable مثل یک برگهٔ سفارش است: روی کاغذ نوشته‌ای «یک پیتزا بپز». این برگه را می‌شود به هر آشپزی داد. اما extends Thread مثل این است که برای هر سفارش یک آشپز مخصوص همان سفارش بسازی که کارش توی وجودش دوخته شده و به درد سفارش بعدی نمی‌خورد. برگهٔ سفارش قابل بازاستفاده است؛ آشپزِ دوخته‌شده نه.

تفاوت Runnable و Callable:

  • Runnable یعنی «بینداز و فراموش کن»: نه مقداری برمی‌گرداند، نه در امضایش اجازهٔ استثنای checked دارد.
  • Callable<V> یعنی «کاری کن و نتیجه‌اش را به من برگردان»: یک مقدار از نوع V برمی‌گرداند و می‌تواند استثنای checked پرتاب کند.
// Runnable: بینداز و فراموش کن، بدون بازگشت، بدون checked در امضا
Runnable r = () -> log.info("done");

// Callable: مقدار برمی‌گرداند، می‌تواند checked پرتاب کند
Callable<Integer> c = () -> {
    return httpClient.fetchCount(); // ممکن است IOException بدهد
};

خب، اگر Callable نتیجه برمی‌گرداند، اما همین حالا اجرا نمی‌شود و بعداً روی نخِ دیگری تمام می‌شود — چطور نتیجه‌اش را بگیرم؟ اینجاست که Future وارد می‌شود.

فیشِ لباس‌شویی

Future<V> مثل فیشی است که خشک‌شویی به تو می‌دهد. کارت (لباس) را سپرده‌ای، الان نتیجه آماده نیست، ولی یک دستگیره داری که با آن بعداً می‌توانی نتیجه را تحویل بگیری. وقتی f.get() را صدا می‌زنی، مثل این است که فیش را ببری و بگویی «حالا لباسم را بده» — و اگر هنوز آماده نباشد، همان‌جا می‌ایستی و منتظر می‌مانی (این یعنی get() مسدودکننده است).

ExecutorService pool = Executors.newFixedThreadPool(4);
Future<Integer> f = pool.submit(c);
try {
    Integer n = f.get(2, TimeUnit.SECONDS); // تا ۲ ثانیه مسدود می‌شود
} catch (ExecutionException e) {
    // وظیفه پرتاب کرد؛ e.getCause() استثنای اصلی است
    Throwable real = e.getCause();
} catch (TimeoutException e) {
    f.cancel(true); // true => نخِ در حال اجرا را interrupt کن
}

سه نکته دربارهٔ Future که مردم را می‌لغزاند — این‌ها را خوب بخوان چون هر سه می‌توانند در تولید باگ‌های خاموش بسازند:

استثناها تا لحظهٔ get() قایم می‌شوند

این خطرناک‌ترین رفتار Future است. اگر وظیفه‌ای که با submit() سپرده‌ای وسط کار استثنا پرتاب کند، هیچ چیز به‌ظاهر اتفاق نمی‌افتد. نه لاگی، نه خطایی، هیچ. استثنا در یک ExecutionException پیچیده می‌شود و فقط و فقط وقتی get() را صدا بزنی به تو تحویل داده می‌شود (و استثنای اصلی در e.getCause() است). در مقابل، اگر با execute(Runnable) — که اصلاً Future نمی‌دهد — کار بسپاری، استثنای گرفته‌نشده به UncaughtExceptionHandler آن نخ می‌رود و معمولاً چاپ می‌شود. این عدم‌تقارن، باگ‌ها را بی‌صدا دفن می‌کند. قاعده: یا همیشه روی Futureهایت get() بزن، یا آگاهانه از execute استفاده کن.

  • f.cancel(true) فقط پرچم وقفه را ست می‌کند؛ خودِ وظیفه باید همکاری کند (پایین‌تر در بخش وقفه توضیحش می‌دهم). cancel(false) جلوی اجرای وظیفه‌ای را که هنوز شروع نشده می‌گیرد، اما وظیفهٔ در حال اجرا را interrupt نمی‌کند.
  • get() مسدودکننده است. یعنی نخِ فراخوان همان‌جا می‌ایستد تا نتیجه آماده شود. اگر می‌خواهی کارها را زنجیر کنی بدون اینکه بایستی، CompletableFuture را ترجیح بده که نامسدودکننده و قابل‌ترکیب است (با متدهایی مثل thenApply, thenCompose, allOf).

چارچوب Executor

حالا پرسش دوم: چه کسی کار را اجرا می‌کند؟ جوابِ حرفه‌ای: یک Executor.

چارچوب Executor یک سلسله‌مراتب تمیز دارد که از ساده به کامل می‌رود:

  • Executor — پایه‌ترین، فقط یک متد دارد: execute(Runnable).
  • ExecutorService — این را گسترش می‌دهد و submit، invokeAll و مدیریتِ چرخهٔ حیات (خاموش‌سازی) را اضافه می‌کند.
  • ScheduledExecutorService — این هم آن را گسترش می‌دهد و تأخیر و اجرای دوره‌ای را اضافه می‌کند.

اما اسبِ‌کارِ واقعی که همهٔ این‌ها پشتش هست، کلاسِ ThreadPoolExecutor است. بیا اول ببینیم چرا نباید کورکورانه به میان‌بُرهای آماده اعتماد کنی.

کارخانهٔ Executors — و چرا باید بهش بدگمان باشی

کلاس Executors چند متدِ کارخانه‌ایِ (factory) راحت دارد که در یک خط برایت استخر می‌سازند. مشکل اینجاست که چندتایشان در محیط تولید خطرناک‌اند، چون یا از صف‌های بی‌کران (unbounded queue) استفاده می‌کنند یا تعداد نخِ بی‌کران:

Executors.newFixedThreadPool(n);   // LinkedBlockingQueue بدون کران ظرفیت → خطر OOM
Executors.newSingleThreadExecutor();// همان صف بی‌کران
Executors.newCachedThreadPool();    // maxPoolSize = Integer.MAX_VALUE → نخ بی‌کران

بگذار خطر را ملموس کنم:

چرا کارخانه‌های آماده در تولید می‌ترکند

«بی‌کران» یعنی «بدون سقف». زیر بارِ مداوم و سنگین، newFixedThreadPool صفِ کارهای منتظرش را بی‌حد رشد می‌دهد — کار روی کار انباشته می‌شود تا هیپ (heap) پر شود و OutOfMemoryError بگیری. از آن طرف، newCachedThreadPool سقفی روی تعداد نخ ندارد (Integer.MAX_VALUE)، پس زیر فشار آن‌قدر نخ می‌سازد تا سیستم‌عامل دیگر نخ نسازد و برنامه از پا دربیاید. رویهٔ سنیور یک جمله است: ThreadPoolExecutor را مستقیم و دستی بساز، با یک صفِ کران‌دار و یک سیاستِ ردّ (rejection policy) صریح. این‌طور وقتی سیستم زیر فشار است، به‌جای ترکیدن، مؤدبانه «نه» می‌گوید.

ThreadPoolExecutor: هفت پارامتر

بیا سازندهٔ کامل را ببینیم. هفت پارامتر دارد و هر کدام یک تصمیم مهندسی است:

ThreadPoolExecutor pool = new ThreadPoolExecutor(
    4,                              // corePoolSize
    16,                             // maximumPoolSize
    60L, TimeUnit.SECONDS,          // keepAliveTime برای نخ‌های بالای core
    new ArrayBlockingQueue<>(1000), // صف کاری کران‌دار
    new ThreadFactoryBuilder()      // نخ‌ها را نام‌گذاری کن! (Guava یا کارخانهٔ دست‌ساز)
        .setNameFormat("api-worker-%d").build(),
    new ThreadPoolExecutor.CallerRunsPolicy() // handlerِ ردّ
);
  • corePoolSize (۴): تعداد نخ‌های همیشگی که حتی وقتی بیکارند نگه داشته می‌شوند.
  • maximumPoolSize (۱۶): بیشترین تعداد نخی که استخر مجاز است تا آن بسازد.
  • keepAliveTime (۶۰ ثانیه): نخ‌های اضافه بر core اگر این‌قدر بیکار بمانند، از بین می‌روند.
  • صف کاری (ArrayBlockingQueue با ظرفیت ۱۰۰۰): جایی که کارها منتظر می‌مانند تا نخی آزاد شود. اینکه این صف کران‌دار باشد، حیاتی است — همین الان می‌فهمی چرا.
  • ThreadFactory: کارخانه‌ای که نخ‌ها را می‌سازد؛ اینجا از آن برای نام‌گذاری نخ‌ها استفاده کرده‌ایم.
  • RejectedExecutionHandler: وقتی همه‌چیز پر است، با کارِ تازه چه کنیم.
الگوریتم ثبت، مثل یک کافه

تصور کن یک کافه داری. corePoolSize یعنی تعداد باریستاهای ثابت (مثلاً ۴). صف، یعنی صندلی‌های انتظار (۱۰۰۰ تا). maximumPoolSize یعنی بیشترین باریستایی که می‌توانی از انبار صدا بزنی (۱۶). حالا وقتی مشتری می‌آید: اول اگر باریستای ثابتِ آزاد کم داری، یکی سرِ کار می‌آید. بعد اگر همه مشغول‌اند، مشتری روی صندلی انتظار می‌نشیند. فقط وقتی همهٔ صندلی‌ها هم پر شد، تازه باریستای اضافه از انبار صدا می‌زنی. و اگر باریستاها هم به سقف رسیدند و صندلی‌ها هم پر بود، دمِ در به مشتری «نه» می‌گویی.

الگوریتم ثبتِ ThreadPoolExecutor دقیقاً همین ترتیب است و ارزش حفظ‌کردن دارد، چون یک شگفتیِ بدنام را توضیح می‌دهد:

  1. اگر نخ‌های در حال اجرا < corePoolSize، نخ جدید بساز برای وظیفه (حتی اگر نخ‌های بیکارِ دیگری وجود داشته باشند).
  2. وگرنه، سعی کن وظیفه را در صف بگذاری.
  3. اگر صف پر است و نخ‌های در حال اجرا < maximumPoolSize، نخ جدید (غیر-core) بساز.
  4. اگر صف پر است و نخ‌ها به max رسیده‌اند، از طریق handler ردّ کن.
شگفتیِ بزرگ: با صف بی‌کران، maximumPoolSize نادیده گرفته می‌شود!

حالا تلّه را ببین. اگر صفت بی‌کران باشد (مثل همان newFixedThreadPool)، گام ۲ همیشه موفق می‌شود — صف که هیچ‌وقت پر نمی‌شود. یعنی شرطِ «صف پر است» در گام ۳ هرگز برقرار نمی‌شود، پس گام ۳ اصلاً فعال نمی‌شود و maximumPoolSize کاملاً بی‌اثر می‌ماند. استخر هرگز از corePoolSize فراتر نمی‌رود، هر چقدر هم بار زیاد شود. این تقریباً همه را غافلگیر می‌کند. پس اگر می‌خواهی استخرت زیر بار بزرگ شود، حتماً باید صفِ کران‌دار بدهی (یا یک SynchronousQueue که ظرفیتش صفر است — پس هر وظیفه یا فوراً یک نخِ نو می‌طلبد یا ردّ می‌شود؛ newCachedThreadPool دقیقاً همین‌طور کار می‌کند).

سیاست‌های ردّ (RejectedExecutionHandler)

وقتی استخر کاملاً اشباع است (صف پر و نخ‌ها در حداکثر)، با کارِ تازه چه کنیم؟ چهار سیاست آماده داری:

سیاست رفتار چه زمانی
AbortPolicy (پیش‌فرض) RejectedExecutionException پرتاب می‌کند می‌خواهی سریع شکست بخوری و فراخوان، سرریز را مدیریت می‌کند
CallerRunsPolicy وظیفه را روی نخِ ثبت‌کننده اجرا می‌کند فشار برگشتی (backpressure): ثبت‌کننده کند می‌شود چون خودش مشغول اجرای وظیفه است، پس ورودی به‌طور طبیعی throttle می‌شود
DiscardPolicy وظیفه را بی‌صدا می‌اندازد دور فقط وقتی از دست دادن کار قابل قبول است (مثلاً متریکِ best-effort)
DiscardOldestPolicy قدیمی‌ترین وظیفهٔ صف را می‌اندازد و ثبت را دوباره امتحان می‌کند به‌ندرت درست؛ کاری را می‌اندازد که بیشترین انتظار را کشیده
فشار برگشتی (backpressure) یعنی چه؟

«فشار برگشتی» یعنی وقتی سیستم زیر فشار است، به‌جای اینکه کورکورانه بیشتر و بیشتر کار قبول کند تا بترکد، به سرچشمه فشار می‌آورد که «آهسته‌تر بفرست». CallerRunsPolicy این را رایگان به تو می‌دهد: وقتی استخر پر است، خودِ نخی که کار را سپرده مجبور می‌شود آن کار را اجرا کند. تا وقتی مشغول اجرای آن است، نمی‌تواند کار جدید بسپارد — پس نرخ ورود خودبه‌خود کم می‌شود. این پیش‌فرضِ عمل‌گرایانه برای خطوط لولهٔ درخواست است. اما یک هشدار: اگر آن نخِ ثبت‌کننده همان نخی باشد که اتصال‌های HTTP را می‌پذیرد، اجرای وظیفه روی آن یعنی تا مدتی پذیرشِ اتصال جدید متوقف می‌شود. که ممکن است خواستهٔ تو باشد یا نباشد.

ریاضیِ اندازه‌گذاری استخر

خب رسیدیم به جایی که نامزدهای سنیور واقعاً امتیاز می‌گیرند: چند نخ بگذارم؟ جواب به یک چیز بستگی دارد — کارت CPU-محور است یا I/O-محور.

حالت CPU-محور: کارِ CPU-محور (فشرده‌سازی، هش، محاسبهٔ محض) یک هسته را تمام‌مدت مشغول نگه می‌دارد. اینجا اگر نخِ بیشتر از تعداد هسته‌ها بسازی، هیچ سودی نمی‌بری؛ فقط سربارِ تعویض‌بافت (context switch — یعنی هزینه‌ای که سیستم‌عامل می‌دهد تا نخی را کنار بگذارد و نخِ دیگری را روی هسته بنشاند) اضافه می‌کنی:

threads ≈ تعداد هسته‌ها        (یا cores + 1 برای پوششِ گاه‌به‌گاهِ page fault)
چرا نخِ بیشتر از هسته سود ندارد

اگر ۴ اجاق داری و ۴ آشپز، هر اجاق یک آشپز، همه‌چیز عالی است. حالا ۸ آشپز بیاور در همان ۴ اجاق. آیا سریع‌تر می‌پزی؟ نه! چون فقط ۴ اجاق هست. آشپزهای اضافه یا بیکار می‌ایستند یا مدام سرِ اجاق با هم جا عوض می‌کنند و همین جابه‌جایی خودش وقت می‌برد. کارِ CPU-محور یعنی «اجاق همیشه روشن است»؛ پس بیشترین سرعت وقتی است که تعداد آشپز = تعداد اجاق.

حالت I/O-محور: کارِ I/O-محور بیشترِ وقتش را منتظر می‌ماند (منتظر دیتابیس، HTTP، دیسک). اینجا داستان کاملاً فرق دارد: می‌خواهی آن‌قدر نخ داشته باشی که وقتی چند نخ منتظر مانده‌اند، بقیه بتوانند از CPU استفاده کنند. فرمول معروفِ برایان گوتز (از کتاب Java Concurrency in Practice)، که در واقع بازگوییِ قانون لیتل (Little's Law) است:

threads = cores × targetUtilization × (1 + waitTime / computeTime)

بیا با یک مثال حل‌شده جا بیندازیم. فرض کن ۸ هسته داری، هدفت ۱۰۰٪ بهره‌وری است، و هر وظیفه ۹۰ میلی‌ثانیه منتظر دیتابیس می‌ماند و ۱۰ میلی‌ثانیه محاسبه می‌کند (پس نسبت انتظار/محاسبه = ۹):

threads = 8 × 1.0 × (1 + 90/10) = 8 × 10 = 80 نخ
نسبت انتظار/محاسبه، مهم‌ترین عددِ اندازه‌گذاری است

دیدی چه شد؟ یک سرویسِ I/O-سنگین به‌درستی حدود ۸۰ نخ روی ۸ هسته می‌خواهد — ده برابرِ تعداد هسته‌ها! منطقش ساده است: چون هر نخ ۹۰٪ وقتش را فقط منتظر است، برای اینکه هسته‌ها بیکار نمانند به نخ‌های زیادی نیاز داری تا همیشه یکی‌شان کار مفید کند. کم بگیری، CPU بیکار می‌ماند و درخواست‌ها صف می‌کشند؛ زیاد بگیری، حافظه می‌سوزانی (هر نخ ~۱ مگابایت پشته) و سیستم را به thrash (جابه‌جاییِ بی‌فایدهٔ مداوم) می‌اندازی. این نسبت انتظار/محاسبه، تک‌عددِ حیاتیِ اندازه‌گذاری استخر است — و دقیقاً همان فشاری است که نخ‌های مجازی (که آخر فصل می‌بینی) برطرفش می‌کنند.

int cores = Runtime.getRuntime().availableProcessors();
double waitToCompute = 9.0; // از پروفایلینگ اندازه‌گیری شده، نه حدس
int size = (int) (cores * (1 + waitToCompute));

یک نکتهٔ عملیِ مهم: عددِ waitToCompute را باید از پروفایلینگِ واقعی دربیاوری، نه از حدس.

همیشه در گلوگاهِ واقعی سقف بگذار

فرمول یک عدد به تو می‌دهد، اما دنیای واقعی یک سقفِ سختِ دیگر هم دارد. اگر استخرِ اتصال دیتابیس‌ات فقط ۲۰ اتصال دارد، بیش از ~۲۰ نخِ همزمانِ دیتابیس فایده‌ای ندارد — نخِ ۲۱ام فقط پشت استخرِ اتصال صف می‌کشد. پس همیشه استخرت را به اندازهٔ گلوگاه (bottleneck) واقعیِ پایین‌دستی کران بگذار. اندازه‌گیری کن که واقعاً کجا تنگ می‌شود و همان‌جا بایست.

ScheduledExecutorService

گاهی نمی‌خواهی کار همین حالا اجرا شود؛ می‌خواهی «۵ ثانیهٔ دیگر» یا «هر ۱۰ ثانیه یک‌بار» اجرا شود. برای این‌ها ScheduledExecutorService را داری.

چرا Timer دیگر مرده است؟

شاید کلاس قدیمیِ Timer را بشناسی. در عمل منسوخ حساب می‌شود، به دو دلیل: (۱) فقط از یک نخ استفاده می‌کند، پس اگر یک وظیفه کند باشد بقیه عقب می‌افتند؛ و (۲) اگر یک وظیفه استثنا پرتاب کند، کلِ Timer بی‌صدا می‌میرد. ScheduledExecutorService جایگزینِ درست و مقاومِ آن است.

ScheduledExecutorService sched = Executors.newScheduledThreadPool(2);

sched.schedule(() -> log.info("once, after 5s"), 5, TimeUnit.SECONDS);

// نرخ ثابت (fixed RATE): هر ۱۰ ثانیه، صرف‌نظر از مدت وظیفه
sched.scheduleAtFixedRate(this::poll, 0, 10, TimeUnit.SECONDS);

// تأخیر ثابت (fixed DELAY): ۱۰ ثانیه بعد از هر اتمام صبر می‌کند — بدون انباشت
sched.scheduleWithFixedDelay(this::poll, 0, 10, TimeUnit.SECONDS);

تفاوت این دو متد را باید در استخوانت حس کنی چون یک تمایز حیاتی است:

قطارِ ساعتی در برابر اتوبوسِ فاصله‌دار

scheduleAtFixedRate مثل قطاری است که ساعتِ حرکتش ثابت است: هر ساعت رأس ساعت راه می‌افتد، مهم نیست سفرِ قبلی چقدر طول کشیده. اگر یک اجرا از دورهٔ ۱۰ ثانیه‌ای بیشتر طول بکشد، اجرای بعدی بلافاصله پشت‌سرش شروع می‌شود تا عقب‌ماندگی را جبران کند (اما هرگز روی یک وظیفهٔ واحد همزمان نمی‌شوند). در مقابل scheduleWithFixedDelay مثل اتوبوسی است که بعد از هر بار برگشتن به پایانه، ۱۰ ثانیه استراحت می‌کند و بعد دوباره راه می‌افتد — یعنی فاصله همیشه بین پایان یکی و شروع بعدی است، پس هرگز انباشت نمی‌شود.

  • scheduleAtFixedRate: یک دورهٔ ثابت را هدف می‌گیرد. اگر اجرا از دوره فراتر رود، اجراهای بعدی پشت‌سرهم برای جبران اجرا می‌شوند.
  • scheduleWithFixedDelay: یک فاصلهٔ ثابت بین پایان یک اجرا و شروع بعدی می‌گذارد، پس اجراها هرگز روی هم نمی‌ریزند.
تلّهٔ مرگِ خاموش

این یکی واقعاً ساعت‌ها از عمرت را می‌تواند بگیرد. اگر یک وظیفهٔ زمان‌بندی‌شده یک استثنای گرفته‌نشده پرتاب کند، آن وظیفه لغو می‌شود و دیگر هرگز اجرا نمی‌شود — استخر آن را بی‌صدا سرکوب می‌کند و هیچ اجرای بعدی‌ای رخ نمی‌دهد، بدون هیچ خطایی در لاگ. تصور کن یک کِران (cron) داری که ماه‌ها کار کرده و یک روز «همین‌طوری» می‌ایستد. علتش تقریباً همیشه همین است. همیشه بدنهٔ وظایف دوره‌ای را در try/catch بپیچ:

sched.scheduleWithFixedDelay(() -> {
    try { doWork(); }
    catch (Throwable t) { log.error("poll failed, will retry next tick", t); }
}, 0, 10, TimeUnit.SECONDS);

خاموش‌سازی: shutdown() در برابر shutdownNow()

یک نکتهٔ مهم: یک ExecutorService را باید خودت خاموش کنی، وگرنه JVM شما اصلاً خارج نمی‌شود! چون نخ‌های کارگرِ استخر به‌طور پیش‌فرض نخِ غیر-daemon (نخِ کاربر) هستند و تا زنده‌اند JVM را زنده نگه می‌دارند. (پایین‌تر daemon را توضیح می‌دهم.)

دو راه خاموش‌سازی داری:

متد وظیفهٔ جدید می‌پذیرد؟ وظایف قبلاً-صف‌شده؟ وظایف در حال اجرا؟
shutdown() نه تا پایان اجرا می‌شوند تمام می‌شوند
shutdownNow() نه به‌صورت List<Runnable> برگردانده می‌شوند، اجرا نمی‌شوند interrupt می‌شوند
بستنِ مؤدبانه در برابر بستنِ فوریِ رستوران

shutdown() یعنی «درِ ورودی را قفل می‌کنیم، مشتری تازه نمی‌پذیریم، اما هر کس داخل است و سفارش داده، غذایش را کامل می‌گیرد». shutdownNow() یعنی «همین الان تعطیل! آشپزها دست از کار بکشید (interrupt)، و سفارش‌هایی که هنوز شروع نشده روی پیشخوان بماند» — آن سفارش‌های شروع‌نشده را به‌صورت یک لیست به تو برمی‌گرداند تا خودت تصمیم بگیری با آن‌ها چه کنی.

توجه کن که shutdownNow() وظایفِ در حال اجرا را interrupt می‌کند (پرچم وقفه‌شان را ست می‌کند)، اما نمی‌تواند وظیفه‌ای را که به وقفه بی‌اعتناست به‌زور متوقف کند. اصطلاحِ متعارف و درستِ خاموش‌سازیِ باوقار این است:

void shutdownGracefully(ExecutorService pool) {
    pool.shutdown(); // پذیرش را قطع کن، بگذار وظایف صف تمام شوند
    try {
        if (!pool.awaitTermination(30, TimeUnit.SECONDS)) {
            pool.shutdownNow(); // به‌زور: وظایف در حال اجرا را interrupt کن
            if (!pool.awaitTermination(10, TimeUnit.SECONDS))
                log.error("pool did not terminate");
        }
    } catch (InterruptedException e) {
        pool.shutdownNow();
        Thread.currentThread().interrupt(); // پرچم را بازیابی کن
    }
}
awaitTermination خودش خاموش نمی‌کند + Java 19 و try-with-resources

دو نکتهٔ ظریف: اول، awaitTermination فقط یک مقدار بولی برمی‌گرداند و خودش خاموش‌سازی را فعال نمی‌کند — باید حتماً اول shutdown() یا shutdownNow() را صدا بزنی، وگرنه تا ابد منتظر می‌مانی. دوم، از جاوا ۱۹ به بعد ExecutorService اینترفیس AutoCloseable را پیاده‌سازی می‌کند؛ متد close() مثل «shutdown() سپس انتظار» رفتار می‌کند، پس می‌توانی استخر را در یک try-with-resources بگذاری و خیالت راحت باشد که خودکار بسته می‌شود.

مدل وقفه (لغو تعاونی)

رسیدیم به پرسش سوم: یک نخ چطور متوقف می‌شود؟ و اینجا جاوا یک تصمیم عمدی و مهم گرفته: هیچ راه امنی برای کشتنِ اجباریِ یک نخ وجود ندارد.

چرا Thread.stop() منسوخ شد

شاید بپرسی «مگر متد Thread.stop() نداریم؟» چرا، ولی منسوخ (deprecated) است و هرگز نباید استفاده‌اش کنی. دلیلش خطرناک است: stop() نخ را وسط کار قطع می‌کند و قفل‌هایش را وسطِ جهش (mutation) — یعنی وسطِ تغییرِ نیمه‌کارهٔ یک شیء — رها می‌کند. نتیجه؟ اشیایی که در حالت خراب و ناسازگار جا می‌مانند. مثل اینکه جراح را وسط عمل به‌زور از اتاق بیرون بکشی. برای همین جاوا این در را بست.

پس لغو در جاوا تعاونی (cooperative) است. یعنی تو فقط درخواست لغو می‌کنی (با interrupt())، و نخِ هدف خودش انتخاب می‌کند که متوجه شود و کارش را تمیز جمع کند.

زدن روی شانهٔ کسی

interrupt() مثل این است که آرام روی شانهٔ یک همکار که هدفون دارد بزنی و بگویی «هر وقت تونستی، کارتو تموم کن و بیا». تو نمی‌توانی مغزش را خاموش کنی؛ فقط یک سیگنال می‌فرستی. اگر او آدمِ همکاری باشد، سرِ اولین فرصت متوجه می‌شود و تمیز جمع می‌کند. اگر بی‌اعتنا باشد و اصلاً چک نکند، هیچ اتفاقی نمی‌افتد. لغوِ تعاونی یعنی همین: همکاریِ نخِ هدف لازم است.

وقفه در واقع فقط یک پرچمِ بولیِ به‌ازای هر نخ است، به‌علاوهٔ چند قاعده:

  • thread.interrupt() — پرچم را ست می‌کند (سیگنال را می‌فرستد).
  • Thread.interrupted()ایستا (static) است؛ پرچمِ نخِ فعلی را می‌خواند و پاک می‌کند.
  • thread.isInterrupted() — پرچم را می‌خواند بدون اینکه پاکش کند.
  • متدهای مسدودکننده‌ای که در امضایشان throws InterruptedException دارند (sleep, wait, join, BlockingQueue.take, Future.get) وقتی interrupt شوند، پرچم را پاک می‌کنند و یک InterruptedException پرتاب می‌کنند.

حالا وقتی یک InterruptedException می‌گیری، دقیقاً دو پاسخِ درست داری:

  1. آن را منتشر کن (بگذار به لایه‌های بالاتر بالا برود)، یا
  2. پرچم را بازیابی کن و از کار خارج شو، تا فراخوان‌های بالای پشته هنوز بتوانند لغو را ببینند:
try {
    queue.take();
} catch (InterruptedException e) {
    Thread.currentThread().interrupt(); // هرگز بی‌صدا نبلع
    return; // واحد کار را متوقف کن
}
گناه کبیره: بلعیدنِ InterruptedException

بدترین کاری که می‌توانی بکنی این است که InterruptedException را با یک catch خالی بگیری و هیچ کاری نکنی. یادت باشد: وقتی این استثنا پرتاب می‌شود، پرچم وقفه پاک شده. حالا اگر تو هم آن را ببلعی، تنها سیگنالی را که به دنیا می‌گفت «بایست» کاملاً نابود کرده‌ای — و حالا وظیفه‌ات دیگر قابل‌لغو نیست و حتی shutdownNow هم نمی‌تواند متوقفش کند. این یکی از پرتکرارترین باگ‌هایی است که در بازبینیِ کدِ سنیور علامت‌گذاری می‌شود. هرگز، هرگز بی‌صدا نبلع.

یک نکتهٔ آخر: متدهای مسدودکننده خودشان InterruptedException می‌دهند، اما اگر یک حلقهٔ CPU-محور داری که هیچ متد مسدودکننده‌ای صدا نمی‌زند، باید پرچم را دستی چک کنی:

while (!Thread.currentThread().isInterrupted()) {
    doOneChunk();
}
// حلقه سرِ interrupt تمیز خارج می‌شود — لغو تعاونی

ThreadLocal — و چرا در استخرها نشت می‌کند

ThreadLocal<T> به هر نخ یک نسخهٔ خصوصیِ خودش از یک متغیر می‌دهد. این مکانیزم استانداردِ نگه‌داشتنِ «زمینهٔ هر درخواست» است — مثل شناسهٔ تراکنش، کاربرِ جاری، یا یک نمونهٔ SimpleDateFormat — بدون اینکه مجبور باشی آن را از میان همهٔ متدها دست‌به‌دست کنی.

کمدِ شخصیِ هر کارمند

ThreadLocal مثل کمدِ شخصیِ هر کارمند در محل کار است. همه به «کمد» دسترسی دارند (متغیر مشترک است)، اما داخلِ کمدِ هر کس فقط مالِ خودش است و کسی محتوای کمدِ دیگری را نمی‌بیند. این عالی است تا وقتی که… کارمندها عوض نشوند. اگر همان کمد را بعداً به یک کارمند تازه بدهی بدون اینکه خالی‌اش کنی، وسایلِ نفرِ قبلی هنوز داخلش است. و دقیقاً همین اتفاق در استخرِ نخ می‌افتد.

private static final ThreadLocal<SimpleDateFormat> FMT =
    ThreadLocal.withInitial(() -> new SimpleDateFormat("yyyy-MM-dd"));

مقادیرِ ThreadLocal در یک نقشه (map) داخلِ خودِ شیءِ Thread ذخیره می‌شوند که کلیدش همان شیءِ ThreadLocal است. این در نخ‌های استخرشده دو خطرِ جدی می‌سازد:

خطر ۱: نشتِ مقدار و دادهٔ کهنه (باگ درستی و امنیت)

نخ‌های استخر عمرِ طولانی دارند و بازاستفاده می‌شوند. مقداری که در حین سرویس‌دادن به درخواستِ A ست کردی، هنوز همان‌جاست وقتی همان نخ بعداً به درخواستِ B سرویس می‌دهد. حالا درخواستِ B زمینهٔ درخواست A را می‌بیند — مثلاً هویتِ کاربرِ A! این هم باگ درستی است هم یک باگِ امنیتی جدی. قاعدهٔ آهنین: اگر در یک نخِ استخرشده یک ThreadLocal را set می‌کنی، باید آن را در یک بلوکِ finally با remove() پاک کنی:

try {
    RequestContext.CURRENT.set(ctx);
    handle(request);
} finally {
    RequestContext.CURRENT.remove(); // در استخرها اجباری
}
خطر ۲: نشتِ حافظه

نقشهٔ درونیِ ThreadLocal (یعنی ThreadLocalMap) برای کلیدهایش (خودِ شیءِ ThreadLocal) از ارجاعِ ضعیف (weak reference) استفاده می‌کند اما برای مقادیر از ارجاعِ قوی (strong reference). حالا فکر کن: اگر خودِ شیءِ ThreadLocal دیگر جایی ارجاع نداشته باشد اما نخ زنده بماند (که نخ‌های استخر دقیقاً همین‌طورند)، مقدار قابل جمع‌آوری توسط زباله‌روب (garbage collector) نیست — تا وقتی آن خانهٔ نقشه دوباره استفاده شود یا نخ بمیرد. اگر اشیای بزرگی را در ThreadLocal کش کرده باشی و استخرت طولانی‌عمر باشد، هیپ آرام‌آرام پر می‌شود. remove() این را هم رفع می‌کند.

InheritableThreadLocal در استخرها کمکت نمی‌کند

شاید فکر کنی InheritableThreadLocal راه‌حل است — این نوع، مقدار را هنگام ساختِ نخِ فرزند به آن کپی می‌کند. اما نخ‌های استخر برای هر وظیفه از نو ساخته نمی‌شوند (نکتهٔ کلِ استخر همین است!)، پس زمینه را به وظایفِ استخر منتقل نمی‌کند. برای انتقالِ درست زمینه در استخرها به بسته‌بندیِ (wrapping) صریح نیاز داری، یا به ScopedValue که در جاوا ۲۱ آمده.

نخ‌های daemon

هر نخ در جاوا یا یک نخِ کاربر (user thread) است یا یک نخِ daemon. قاعدهٔ کلیدی: JVM دقیقاً وقتی خارج می‌شود که آخرین نخِ غیر-daemon تمام شود — و در آن لحظه، هر نخِ daemonی که هنوز در حال کار باشد را بی‌رحمانه رها می‌کند (بدون هیچ تضمینی که بلوک‌های finally اجرا شوند).

کارمند رسمی در برابر نگهبانِ شب

نخِ کاربر مثل کارمند رسمی است: تا او کاری دارد، ساختمان (JVM) باز می‌ماند. نخِ daemon مثل نگهبانِ شب یا سیستمِ تهویه است: کارِ پس‌زمینه انجام می‌دهد، اما لحظه‌ای که آخرین کارمندِ رسمی برود، چراغ‌ها خاموش می‌شود و نگهبان هم — چه بخواهد چه نخواهد، حتی اگر وسطِ کاری باشد — بیرون می‌ماند. برای همین کارِ حیاتی را هرگز به نگهبانِ شب نمی‌سپاری.

Thread t = new Thread(task);
t.setDaemon(true);   // باید پیش از start() ست شود، وگرنه IllegalThreadStateException
t.start();
  • از نخِ daemon برای زیرساختِ پس‌زمینه استفاده کن که هرگز نباید برنامه را زنده نگه دارد: فلاش‌کردنِ متریک‌ها، تایمرهای پاک‌سازیِ کش، نظرسنجی‌های ضربان (heartbeat).
  • هرگز از نخِ daemon برای کاری که باید تمام شود یا باید پاک‌سازی کند استفاده نکن (نوشتن یک فایل، فلاش‌کردن یک تراکنش) — چون JVM ممکن است وسطِ نوشتن آن را بکشد.
  • یک نخ به‌طور پیش‌فرض وضعیتِ daemonِ والدش را به ارث می‌برد. برای همین است که نخ‌های Executors نخِ کاربر (غیر-daemon) هستند و تا وقتی استخر را خاموش نکنی JVM را زنده نگه می‌دارند. اگر نخ‌های daemon در استخر می‌خواهی، باید یک ThreadFactory بدهی که setDaemon(true) را صدا بزند.

نخ‌های مجازی (Java 21+) — جایی که ریاضیِ اندازه‌گذاری بخار می‌شود

و حالا بزرگ‌ترین تغییرِ سال‌های اخیر. جاوا ۲۱ نخ‌های مجازی (virtual threads) را دائمی کرد. بگذار اول شهودش را بسازم.

پیک‌های موتوری و ماشین‌های شرکت

تصور کن یک شرکتِ پیک داری. نخِ سکویی (platform thread) مثل یک ماشینِ گران‌قیمتِ شرکت است: تعدادش محدود است و هر کدام گران. اگر پیکی برود دمِ در و منتظرِ باز شدنِ در بماند، آن ماشینِ گران همان‌جا بیکار پارک است. نخِ مجازی مثل یک پیکِ کم‌هزینه است: وقتی به دری می‌رسد و باید منتظر بماند، از ماشین پیاده می‌شود (unmount) و ماشین را آزاد می‌کند تا پیکِ دیگری سوارش شود؛ وقتی در باز شد، دوباره سوارِ یک ماشینِ آزاد می‌شود. تعداد ماشین‌ها کم است (نخ‌های حاملِ سکویی = carrier)، اما تعداد پیک‌ها می‌تواند میلیون‌ها باشد.

پس نخِ مجازی نخی است که خودِ JVM آن را روی یک استخرِ کوچک از نخ‌های حاملِ (carrier) سکویی زمان‌بندی می‌کند؛ وقتی روی I/O مسدود می‌شود، به‌جای اینکه یک نخِ سیستم‌عامل را مسدود کند، پیاده می‌شود (unmount) — در فضای کاربر پارک می‌شود و حاملش را آزاد می‌کند. آن‌قدر ارزان‌اند که می‌توانی میلیون‌ها بسازی.

// هر وظیفه نخ مجازی خودش را می‌گیرد؛ این یک استخر نیست که اندازه‌اش کنی
try (var exec = Executors.newVirtualThreadPerTaskExecutor()) {
    for (var req : requests)
        exec.submit(() -> handleBlocking(req)); // فراخوان‌های مسدودکننده اینجا اشکالی ندارند
} // close() منتظر همهٔ وظایف می‌ماند (AutoCloseable)
تغییرِ پارادایم: دیگر استخر را اندازه نمی‌گیری

این جانِ کلام است. برای کارِ I/O-محور، دیگر استخر را اندازه‌گذاری نمی‌کنی. کدِ سرراستِ مسدودکننده می‌نویسی، یک نخِ مجازی به‌ازای هر وظیفه، و می‌گذاری JVM آن‌ها را روی حامل‌های کم مالتی‌پلکس (چندگانه‌سازی) کند. یادت هست فرمول گوتز؟ آن فرمول در واقع یک راهِ دور زدنِ مشکلِ «نخِ سکوییِ گران» بود. نخِ مجازی همان محدودیتی را که فرمول برایش جبران می‌کرد، از ریشه حذف می‌کند.

اما چند مورد که هنوز می‌گزند:

پین‌شدن (pinning) — قاتلِ خاموشِ نخ مجازی

اگر یک نخِ مجازی داخلِ یک بلوکِ synchronized (یا وسطِ یک فراخوانِ بومی/JNI) باشد و همان‌جا مسدود شود، دیگر نمی‌تواند از حاملش پیاده شود — به آن حامل پین (pin) می‌شود، یعنی به آن میخکوب می‌ماند. اگر به‌قدر کافی نخ به حامل‌ها پین شوند، استخرِ حامل‌ها گرسنه می‌شود و توان عملیاتی می‌خوابد. راه‌حل: حولِ I/Oِ مسدودکننده به‌جای synchronized از ReentrantLock استفاده کن که به نخِ مجازی اجازهٔ پیاده‌شدن می‌دهد. (نسخه‌های بعدیِ JDK، یعنی ۲۴ به بعد، بیشترِ پین‌شدنِ ناشی از synchronized را حذف کرده‌اند، اما روی خودِ جاوا ۲۱ این مشکل واقعی است.)

  • استخرشان نکن. ساختنِ newFixedThreadPool از نخ‌های مجازی کلِ فلسفه را نقض می‌کند؛ همیشه به‌ازای هر وظیفه یکی بساز.
  • برای کارِ CPU-محور هیچ سودی ندارند. نخِ مجازی به انتظار کمک می‌کند نه به محاسبه. برای محاسبهٔ محض همان استخرِ نخِ سکوییِ اندازه‌گرفته‌شده به هسته‌ها را نگه دار.

تلّه‌ها و نکته‌های رایج (رگبار سریع)

بیا همهٔ تلّه‌هایی را که در فصل دیدیم یک‌جا مرور کنیم — این‌ها همان‌هایی‌اند که در تولید خون‌دل درست می‌کنند:

  • صدا زدنِ run() به‌جای start() — بدون نخ جدید.
  • صفِ بی‌کرانِ newFixedThreadPool → OOM زیر بار.
  • maximumPoolSize نادیده گرفته می‌شود چون صف بی‌کران است.
  • بلعیدنِ InterruptedException → وظایفِ غیرقابل‌لغو.
  • فراموش‌کردنِ ThreadLocal.remove() در استخر → زمینهٔ نشت‌کرده/کهنه.
  • نام‌گذاری‌نکردنِ نخ‌ها → thread dumpهای بی‌فایده در تولید. همیشه ThreadFactory بده.
  • استثناهای وظایفِ submitشده تا get() بی‌صدا گم می‌شوند.
  • وظیفهٔ زمان‌بندی‌شده یک‌بار پرتاب کند → برای همیشه بی‌صدا می‌ایستد.
  • اشتراکِ شیءِ غیرِ thread-safe (SimpleDateFormat، Random زیر رقابت) بین نخ‌های استخر.

بهترین روش‌ها

۱. ThreadPoolExecutor را صریح با صف کران‌دار + سیاست ردّ انتخابی بساز؛ در تولید از کارخانه‌های بی‌کرانِ Executors بپرهیز. ۲. نخ‌هایت را نام‌گذاری کن از طریق ThreadFactory — خودِ آینده‌ات هنگام خواندنِ thread dump سپاسگزار می‌شود. ۳. استخرهای CPU را به cores اندازه بگیر؛ استخرهای I/O را با فرمولِ انتظار/محاسبه، سپس در گلوگاهِ واقعیِ پایین‌دستی سقف بگذار. ۴. همیشه هنگام پایین‌آمدن shutdown() + awaitTermination() + shutdownNow(). ۵. با InterruptedException مثل سیگنالی درجه‌یک برخورد کن: منتشر کن یا بازیابی-و-خروج، هرگز نبلع. ۶. در استخرها هر ThreadLocal را در finally با remove() پاک کن. ۷. CompletableFuture/همزمانیِ ساختاریافته را بر زنجیره‌های get()ِ مسدودکننده ترجیح بده. ۸. در جاوا ۲۱ به بعد، برای fan-outِ I/O-محور از نخ‌های مجازی استفاده کن و تنظیمِ دستیِ اندازهٔ استخر را کنار بگذار.

پرسش‌های مصاحبه

حالا بیا هر چه یاد گرفتی را در قالبِ پرسش‌های واقعیِ مصاحبه تمرین کنیم. هر کدام را اول خودت جواب بده، بعد پاسخ را بخوان.

۱. تفاوت RUNNABLE و BLOCKED چیست؟ چرا خواندن از سوکت RUNNABLE است؟

BLOCKED یعنی نخ مشخصاً بر سرِ قفلِ مانیتورِ synchronized رقابت می‌کند. WAITING/TIMED_WAITING یعنی park شده و منتظرِ سیگنال است. RUNNABLE یعنی JVM آن را قادر به اجرا می‌داند. JVM انسدادِ سطحِ سیستم‌عامل روی I/O را نمی‌بیند، پس نخی که در خواندنِ مسدودکنندهٔ سوکت گیر کرده هنوز RUNNABLE برچسب می‌خورد — انسداد زیرِ دیدِ JVM رخ می‌دهد.

۲. اگر start() را دوبار صدا بزنی چه می‌شود؟ صدا زدنِ مستقیم run() چطور؟

start() دوم IllegalThreadStateException پرتاب می‌کند — نخ یک‌بارمصرف است. صدا زدنِ run() بدنه را هم‌ریسمانی روی نخِ فعلی اجرا می‌کند؛ نخِ جدیدی ساخته نمی‌شود. این باگی تصادفیِ رایج است.

۳. [سخت] یک ThreadPoolExecutor با core=۴، max=۵۰ و صفِ بی‌کرانِ سبکِ newFixedThreadPool می‌سازی. زیر بار سنگین چند نخ اجرا می‌شود؟

فقط ۴. با صفِ بی‌کران، وظایف همیشه با موفقیت صف می‌شوند، پس استخر هرگز به شرطِ «صف پر» که ساختِ نخِ بالای core را فعال می‌کند نمی‌رسد. maximumPoolSize عملاً مرده است. برای استفاده از max به صفِ کران‌دار نیاز داری.

۴. الگوریتم دقیقِ ثبتِ ThreadPoolExecutor را مرور کن.

(۱) اگر نخ‌ها < core، نخ جدید بساز. (۲) وگرنه سعی کن صف کنی. (۳) اگر صف‌کردن شکست خورد (صف پر) و نخ‌ها < max، نخِ غیر-core بساز. (۴) وگرنه handlerِ ردّ را فراخوان کن. نکتهٔ ظریف: صف‌کردن پیش از رشد فراتر از core امتحان می‌شود، به همین دلیل صفِ بی‌کران شما را در core سقف می‌کند.

۵. چهار سیاست ردّ را مقایسه کن و بگو کدام فشار برگشتی می‌دهد.

AbortPolicy پرتاب می‌کند (fail-fast، پیش‌فرض). DiscardPolicy بی‌صدا می‌اندازد. DiscardOldestPolicy سرِ صف را می‌اندازد و دوباره امتحان می‌کند. CallerRunsPolicy وظیفه را روی نخِ ثبت‌کننده اجرا می‌کند — این فشار برگشتی می‌دهد چون ثبت‌کننده مشغول است، نمی‌تواند بیشتر ثبت کند و ورودی را throttle می‌کند.

۶. تعداد نخ را برای این استخراج کن: ۱۶ هسته، وظایفی که ۱۹۰ms روی یک سرویس پایین‌دستی منتظرند و ۱۰ms محاسبه می‌کنند، هدف ۱۰۰٪ بهره‌وری.

threads = cores × util × (1 + wait/compute) = 16 × 1.0 × (1 + 190/10) = 16 × 20 = 320. اما در سقفِ محدودیتِ همزمانیِ واقعیِ پایین‌دستی هم کران بگذار — ۳۲۰ نخ که به سرویسی با ۵۰ همزمانِ مجاز می‌کوبند فقط صف می‌کشند؛ به اندازهٔ گلوگاه بگیر.

۷. [سخت] چرا بلعیدنِ InterruptedException باگ جدی است و دو مدیریتِ درست کدام‌اند؟

پرچمِ وقفه هنگامِ پرتابِ استثنا پاک می‌شود. اگر آن را بگیری و کاری نکنی، تنها سیگنالِ لغو را نابود کرده‌ای — وظیفه غیرقابل‌interrupt می‌شود و shutdownNow نمی‌تواند متوقفش کند. درست: یا دوباره پرتاب/منتشرش کن، یا Thread.currentThread().interrupt() را صدا بزن تا پرچم بازیابی شود و سپس از کار خارج شو.

۸. shutdown() در برابر shutdownNow() — در هرکدام چه بر سرِ وظایفِ صف‌شده و در حال اجرا می‌آید؟

shutdown(): پذیرشِ وظیفهٔ جدید را قطع می‌کند، می‌گذارد وظایفِ صف‌شده و در حال اجرا تمام شوند. shutdownNow(): پذیرش را قطع می‌کند، وظایفِ صف‌شدهٔ شروع‌نشده را به‌صورت List<Runnable> برمی‌گرداند (هرگز اجرا نمی‌شوند) و وظایفِ در حال اجرا را interrupt می‌کند (که فقط اگر به وقفه احترام بگذارند متوقفشان می‌کند).

۹. [سخت] چرا استفاده از ThreadLocal در یک استخر نخ هم باگ درستی می‌سازد هم نشت حافظه؟

درستی: نخ‌های استخر بازاستفاده می‌شوند، پس مقداری که برای یک وظیفه ست شده به وظیفهٔ بعدی روی همان نخ نشت می‌کند مگر remove() کنی — زمینهٔ یک درخواست (مثلاً هویتِ کاربر) به دیگری نشت می‌کند. حافظه: کلیدهای ThreadLocalMap ضعیف‌اند اما مقادیر قوی؛ نخِ طولانی‌عمرِ استخر مقدار را نگه می‌دارد حتی بعد از رفتنِ ThreadLocal، تا خانه بازاستفاده شود. remove() در finally هر دو را رفع می‌کند.

۱۰. این چه چاپ می‌کند؟
ExecutorService pool = Executors.newSingleThreadExecutor();
Future<?> f = pool.submit(() -> { throw new RuntimeException("boom"); });
System.out.println("submitted");
// بدون فراخوان f.get()
pool.shutdown();

submitted را چاپ می‌کند و هیچ‌چیز دربارهٔ «boom». استثناهای وظایفِ submitشده در Future گرفته می‌شوند و فقط روی f.get() رو می‌شوند (به‌صورت ExecutionException). بدون get()، استثنا بی‌صدا بلعیده می‌شود. (اگر از execute() استفاده کرده بودی، به UncaughtExceptionHandler می‌رسید و stack trace چاپ می‌کرد.)

۱۱. scheduleAtFixedRate در برابر scheduleWithFixedDelay وقتی یک اجرا از دوره‌اش فراتر می‌رود؟

scheduleAtFixedRate یک دورهٔ ثابت را هدف می‌گیرد؛ اگر اجرا از دوره فراتر رود، اجرای بعدی بلافاصله بعدش شروع می‌شود (برای جبران burst می‌کند، هرگز روی وظیفهٔ واحد همپوشانی نمی‌کند). scheduleWithFixedDelay همیشه تأخیرِ کامل را بعد از اتمام می‌گذارد، پس اجراها هرگز انباشته نمی‌شوند. برای نظرسنجی‌ای که نباید همپوشانی یا هجوم کند، از تأخیرِ ثابت استفاده کن.

۱۲. [سخت] یک وظیفهٔ scheduleWithFixedDelay بعد از مدتی بدون خطا در لاگ می‌ایستد. چرا؟

استثنای گرفته‌نشده پرتاب کرد. وظیفهٔ زمان‌بندی‌شده‌ای که پرتاب کند بی‌صدا لغو می‌شود و دیگر هرگز زمان‌بندی نمی‌شود؛ استثنا در Futureِ (نادیده‌گرفته‌شده) دام می‌افتد. بدنهٔ وظیفه را در try/catch بپیچ تا زنده بماند.

۱۳. چرا setDaemon(true) باید پیش از start() صدا زده شود و چه زمانی یک نخ daemon داده از دست می‌دهد؟

وضعیتِ daemon وقتی نخ زنده شد تغییرناپذیر است — ست‌کردنش بعد از start() استثنای IllegalThreadStateException می‌دهد. یک daemon وقتی داده از دست می‌دهد که تنها اجراکنندهٔ عملیاتی حیاتی باشد و همهٔ نخ‌های کاربر خارج شوند: JVM خاتمه می‌یابد و daemon را وسطِ عملیات رها می‌کند، بدون تضمینِ اجرای finally/پاک‌سازی. هرگز کارِ باید-تکمیل‌شود را روی daemonها نگذار.

۱۴. [سخت] در جاوا ۲۱ یک سرویس را به نخ‌های مجازی می‌بری اما توان عملیاتی بهتر نمی‌شود و CPU نشان می‌دهد نخ‌های حامل بیشتر بیکارند. علت محتمل چیست؟

پین‌شدن (pinning). I/Oِ مسدودکننده درونِ بلوکِ synchronized (یا فراخوانِ بومی) نشسته، پس نخِ مجازی هنگامِ مسدودشدن نمی‌تواند از حاملش پیاده شود. با حامل‌های کم، نخ‌های مجازیِ پین‌شده استخر را گرسنه می‌کنند. رفع: synchronized حولِ فراخوانِ مسدودکننده را با ReentrantLock جایگزین کن که به نخِ مجازی اجازهٔ پیاده‌شدن می‌دهد. (همچنین بررسی کن نخ‌های مجازیِ به‌ازای-وظیفه استفاده می‌کنی، نه استخرشده.)

۱۵. چه زمانی نخ‌های مجازی ابزار نادرست‌اند؟

برای کارِ CPU-محور — موازی‌سازی‌ای فراتر از تعدادِ حامل‌ها که به هسته‌ها کران دارد اضافه نمی‌کنند. وظایفِ محاسبه‌سنگین باید از استخرِ نخِ سکوییِ اندازه‌گرفته‌شده به هسته‌ها استفاده کنند. نخ‌های مجازی فقط وقتی می‌برند که نخ‌ها وقتشان را منتظر بگذرانند.

نکاتِ سنیور و موارد پیشرفته

تا اینجا مدل ذهنی نخ، Executor، اندازه‌گیری pool، لغو تعاونی و virtual threads را داری. حالا می‌رویم سراغ لایه‌ای که فقط بعد از چند شبِ on-call و چند post-mortem واقعی یاد می‌گیری: چیزهایی که در کد درست به نظر می‌رسند ولی زیر بار prod می‌ترکند. این بخش هیچ‌کدام از مطالب قبلی را تکرار نمی‌کند؛ فقط شکاف‌ها را پر می‌کند.

نقشهٔ این بخش

۱. مدل حافظه (JMM): چه garantee‌ای مجانی از Executor می‌گیری و چرا برای خواندن نتیجهٔ task نیازی به volatile نداری. ۲. availableProcessors() در container دروغ می‌گوید — بزرگ‌ترین اشتباه sizing در Kubernetes. ۳. ForkJoinPool.commonPool() و parallel stream یک pool مشترک‌اند و blocking داخلشان کل JVM را قحطی‌زده می‌کند. ۴. CompletableFuture درست — کدام thread callback را اجرا می‌کند، join در برابر get، جمع‌کردن نتایج. ۵. Thread-starvation deadlock (pool تو در pool). ۶. Observability و hookها، propagation کانتکست/MDC، انواع queue. ۷. دام‌های @Async در Spring. ۸. پلی‌بوک واقعی virtual threads: Semaphore جای sizing، pinning در JDK جدید، ScopedValue و structured concurrency. ۹. هفت سؤال سختِ سنیور با پاسخ کامل.

مدل حافظه: happens-before‌ای که مجانی می‌گیری

یک سؤال ظریف: task روی نخ B مقداری را حساب می‌کند و در یک فیلد معمولی (بدون volatile) می‌نویسد؛ نخ A بعد از future.get() آن فیلد را می‌خواند. آیا A حتماً مقدار تازه را می‌بیند یا ممکن است مقدار کهنه/نیمه‌ساخته ببیند؟

ضمانت‌های JMM برای Executor

مستندات پکیج java.util.concurrent دو رابطهٔ happens-before را صریحاً تضمین می‌کنند: ۱. هر چیزی که قبل از submit/execute روی نخ فرستنده انجام دادی، happens-before شروع اجرای task است. یعنی task همهٔ نوشته‌های قبل از submit را می‌بیند. ۲. هر چیزی که task انجام می‌دهد، happens-before بازگشتِ Future.get() است. یعنی بعد از get() نخ A تضمیناً همهٔ نوشته‌های task را به‌روز می‌بیند. پس نیازی به volatile یا synchronized برای انتقال نتیجهٔ task نداری؛ خودِ مرز submit→run→get این visibility را می‌سازد. همین ضمانت برای BlockingQueue (put/take)، CountDownLatch، و اتمام یک نخ نسبت به join() هم برقرار است.

دام: fire-and-forget بدون هیچ نقطهٔ همگام‌سازی

اگر task را با execute() بفرستی و نتیجه را بدون هیچ سازوکار همگام‌سازی (نه get، نه queue، نه latch، نه فیلد volatile) از نخ دیگری بخوانی، هیچ رابطهٔ happens-before نداری و ممکن است برای همیشه مقدار کهنه یا حتی object نیمه‌ساخته ببینی. باگش هم غیرقطعی است و در تست دیده نمی‌شود.

availableProcessors() در container دروغ می‌گوید

فرمول‌های sizing فصل قبل همه به «تعداد coreها» تکیه دارند. اما آن عدد از کجا می‌آید؟ از Runtime.getRuntime().availableProcessors(). مشکل اینجاست که در Kubernetes این عدد اغلب همان چیزی که فکر می‌کنی نیست.

بزرگ‌ترین اشتباه sizing در سال‌های اخیر

node تو ۳۲ core دارد، ولی به pod یک cpu limit مثلاً 500m (نصف یک core) داده‌ای. JVM از JDK 10 به بعد container-aware است و cgroup را می‌خواند؛ با quota زیر یک core، availableProcessors() مقدار ۱ برمی‌گرداند. حالا سه فاجعه:

  • اگر pool را با عددِ ثابت ۳۲ ساخته بودی، ۳۲ نخ روی سهمیهٔ یک core دائم context-switch می‌کنند و throttle می‌شوی.
  • ForkJoinPool.commonPool() که parallelism‌اش availableProcessors()-1 است، عملاً ۰ (یعنی یک نخ) می‌شود و همهٔ parallel streamها سریال اجرا می‌شوند.
  • برعکس، اگر limit نگذاری ولی JVM قدیمی/غلط پیکربندی باشد، ممکن است ۳۲ ببیند و روی همهٔ همسایه‌ها اثر بگذارد.
چه‌کار کنیم
  • cgroups v2 از JDK 17 به بعد کامل پشتیبانی می‌شود؛ روی JDKهای قدیمی و k8s جدید حواست باشد.
  • اگر limit را کسری گذاشتی ولی می‌خواهی JVM عدد مشخصی ببیند، -XX:ActiveProcessorCount=N را صریح ست کن.
  • عدد واقعی را در startup لاگ کن: log.info("cores={}", Runtime.getRuntime().availableProcessors()). همین یک خط ساعت‌ها debugging را حذف می‌کند.
  • ترجیحاً cpu request == cpu limit بگذار تا رفتار قابل‌پیش‌بینی شود.

commonPool() و parallel streamها یک pool مشترک‌اند

stream().parallel() و CompletableFuture.supplyAsync(task) (نسخهٔ بدون executor) هر دو روی یک pool سراسری اجرا می‌شوند: ForkJoinPool.commonPool(). این pool کوچک است (به اندازهٔ coreها) و برای کار CPU-bound کوتاه ساخته شده.

قحطی commonPool: یک blocking، کل برنامه را زمین می‌زند

اگر داخل یک parallelStream().map(...) یک فراخوانی blocking (HTTP، DB، Thread.sleep) بگذاری، آن worker نخِ commonPool را قفل می‌کند. چون این pool در کل JVM مشترک است، حالا هر بخش دیگری از برنامه که به parallel stream یا supplyAsync تکیه دارد گرسنه می‌ماند و درخواست‌های بی‌ربط کند می‌شوند. کلاسیک‌ترین حادثه: یک endpoint گزارش‌گیری با parallel stream، کل سرویس را با هم می‌خواباند.

دو راه‌حل درست:

// راه ۱: parallel stream را داخل یک ForkJoinPool اختصاصی اجرا کن
ForkJoinPool custom = new ForkJoinPool(16);
List<R> out = custom.submit(() ->
    items.parallelStream().map(this::blockingCall).toList()
).get();

// راه ۲ (برای CompletableFuture): همیشه executor خودت را پاس بده
CompletableFuture.supplyAsync(this::blockingCall, myIoPool);
ManagedBlocker: راه اصولی برای blocking داخل ForkJoinPool

اگر مجبوری داخل FJP block کنی، به‌جای block ساده از ForkJoinPool.managedBlock(...) استفاده کن. این interface به pool خبر می‌دهد که «من دارم block می‌شوم»، و pool موقتاً یک worker جبرانی راه می‌اندازد تا parallelism حفظ شود. CompletableFuture هم داخلش از همین ManagedBlocker برای join() روی worker نخ‌های FJP استفاده می‌کند.

CompletableFuture درست: کجا callback اجرا می‌شود

فصل قبل CompletableFuture را در حد «non-blocking و composable» معرفی کرد. اما نکتهٔ سنیوری این است: callback کجا اجرا می‌شود؟

  • نسخهٔ بدون Async (مثل thenApply): callback روی نخی اجرا می‌شود که stage قبلی را کامل کرده — یا حتی روی نخِ فراخواننده اگر آن لحظه از قبل complete شده باشد. غیرقابل‌پیش‌بینی و خطرناک برای کارِ سنگین.
  • نسخهٔ Async بدون executor (thenApplyAsync): روی commonPool() — با همان دام قحطی بالا.
  • نسخهٔ Async با executor: روی pool خودت. این حالت پیش‌فرضِ درستِ prod است.
CompletableFuture
    .supplyAsync(() -> fetchUser(id), ioPool)     // I/O روی pool ما
    .thenApplyAsync(this::enrich, cpuPool)         // CPU روی pool جدا
    .exceptionally(ex -> User.anonymous())         // fallback
    .thenAccept(this::send);
`thenApply` در برابر `thenCompose`، `join` در برابر `get`، و مدیریت خطا

اگر تابعت خودش یک CompletableFuture برمی‌گرداند، thenApply به CompletableFuture<CompletableFuture<T>> تودرتو می‌رسد؛ برای flatten از thenCompose (معادل flatMap) استفاده کن. get() استثنای checked می‌دهد؛ join() نسخهٔ unchecked است (CompletionException) و داخل زنجیره تمیزتر. برای جمع‌کردن چند future: allOf(a,b,c).thenApply(...) — ولی allOf مقدار Void برمی‌گرداند، پس بعدش تک‌تک join بزن. برای خطا: exceptionally فقط مسیر خطا، handle هر دو مسیر، و whenComplete یک finally است که مشاهده می‌کند ولی تغییر نمی‌دهد.

Thread-starvation deadlock: pool تو در pool

این یکی از موذی‌ترین deadlockهاست و در تست هیچ‌وقت دیده نمی‌شود چون فقط زیر بار رخ می‌دهد.

توالی: pool تو کراندار است (مثلاً ۱۰ نخ). یک task روی این pool، خودش subtaskهایی را به همان pool submit می‌کند و بعد get() می‌زند و منتظر نتیجهٔ آن‌ها می‌ماند. اگر ۱۰ taskِ والد همزمان همهٔ ۱۰ نخ را بگیرند و هرکدام منتظر subtask خودش باشد، دیگر هیچ نخی برای اجرای subtaskها آزاد نیست — همه منتظرِ همه؛ قفل کامل.

نمودار — چرا pool تودرتوی کراندار قفل می‌شود:

sequenceDiagram
    participant P as Parent tasks (fill all N threads)
    participant Q as Work queue
    participant C as Child tasks
    P->>Q: submit child, then block on get()
    Q-->>C: waiting for a free thread
    Note over P,C: all N threads held by parents<br/>no thread free to run children
    Note over P,C: parents wait children, children wait threads = deadlock
قانون طلایی

یک taskی که روی یک pool اجرا می‌شود نباید به همان pool کار submit کند و روی نتیجه‌اش block شود. یا poolهای جدا برای لایه‌های مختلف بگذار، یا کل زنجیره را non-blocking با thenCompose بنویس، یا از virtual threads استفاده کن (که نخ کمیاب نیست و این کلاس deadlock را حذف می‌کند).

Observability و hookها

poolی که نمی‌توانی ببینی یک روز غافلگیرت می‌کند. getActiveCount()، getQueue().size() و getCompletedTaskCount() را به Micrometer بده — صفِ همیشه‌پر یعنی pool کوچک یا downstream کند. و beforeExecute/afterExecute را در ThreadPoolExecutor override کن: جای عالی برای ست/پاک‌کردن MDC، اندازه‌گیری زمان اجرا، و گرفتنِ استثناهایی که در submit گم می‌شوند (afterExecute هم Runnable و هم Throwable را می‌گیرد). دو دستهٔ دیگر: allowCoreThreadTimeOut(true) اجازه می‌دهد نخ‌های coreِ بی‌کار هم جمع شوند (poolهای کم‌مصرف)، و prestartAllCoreThreads() نخ‌های coreِ lazy را از قبل warm می‌کند وقتی latencyِ اولین درخواست‌ها مهم است.

propagation کانتکست/MDC در async

ThreadLocal‌های SLF4J MDC، SecurityContextHolder و trace-idهای tracing خودکار به نخ pool منتقل نمی‌شوند؛ نخ کارگر آن‌ها را ندارد. راه‌حل: یک TaskDecorator (در Spring) یا wrapper که قبل از اجرا کانتکست را از نخ فرستنده کپی و بعدش پاک کند. با virtual threads این درد کمتر می‌شود ولی حذف نمی‌شود.

انتخاب نوع queue هم یک تصمیم مهندسی است: ArrayBlockingQueue پیش‌فرضِ امن prod است، PriorityBlockingQueue برای صف با اولویت، و DelayQueue برای زمان‌بندی — همه به‌جای LinkedBlockingQueueی بی‌کران و خطرناک.

دام‌های @Async در Spring

@Async قشنگ‌ترین «برو async» دنیاست تا وقتی یکی از این‌ها بزندت:

سه دام کلاسیک `@Async`

۱. self-invocation: @Async با یک proxy کار می‌کند. اگر یک متد public از همان bean متد @Async را صدا بزند، فراخوانی از proxy رد نمی‌شود و کاملاً سنکرون اجرا می‌شود — بدون هیچ خطا. باید از یک bean دیگر صدایش بزنی. ۲. گم‌شدن استثنا: اگر متد @Async بازگشت void داشته باشد و خطا بدهد، استثنا هیچ‌جا نمی‌رود (مثل execute بدون Future) و در لاگ گم می‌شود؛ برای گرفتنش باید AsyncUncaughtExceptionHandler تعریف کنی، یا بازگشت را CompletableFuture کنی تا خطا در get دربیاید. ۳. executor پیش‌فرض: در Spring خالص (بدون Boot) اگر executor مشخص نکنی، fallback به SimpleAsyncTaskExecutor است که pool ندارد و برای هر فراخوانی یک نخ نو می‌سازد — زیر بار همان انفجار نخ. در Spring Boot خوشبختانه یک ThreadPoolTaskExecutor خودکار می‌سازد، ولی به آن هم default اعتماد نکن و صریح پیکربندی کن.

Spring Boot و virtual threads (نسخهٔ مدرن)

از Spring Boot 3.2 به بعد با spring.threads.virtual.enabled=true (روی Java 21+) هم executor @Async و هم نخ‌های Tomcat به virtual threads سوئیچ می‌شوند و کدِ blocking سنتی‌ات بدون بازنویسی مقیاس می‌گیرد. ولی همان هشدارهای MDC و pinning همچنان برقرارند.

پلی‌بوک واقعی virtual threads

فصل قبل گفت با virtual threads دیگر pool را sizing نمی‌کنی. سؤال بعدیِ سنیور: پس جلوی downstream محدود را چطور بگیرم؟

Semaphore جایگزینِ «اندازهٔ pool» است

وقتی یک virtual thread برای هر task می‌سازی، دیگر pool مرزی ندارد؛ ولی DB تو شاید فقط ۲۰ connection بدهد. اگر ۱۰٬۰۰۰ virtual thread همزمان به آن حمله کنند، connection pool منفجر یا timeout می‌شود. راه‌حل: concurrency را نه با اندازهٔ pool، بلکه با یک Semaphore به منبعِ کمیاب bound کن:

Semaphore db = new Semaphore(20);
db.acquire();
try { return jdbc.query(...); }
finally { db.release(); }

در دنیای virtual threads، Semaphore همان نقشی را دارد که قبلاً maximumPoolSize داشت: محافظ ظرفیت downstream.

ThreadLocal و شیءهای گران در دنیای میلیونی

با platform threadها یک ترفند رایج، cacheِ یک شیء گران (مثلاً بافر بزرگ) در ThreadLocal بود؛ چون نخ‌ها کم بودند. با میلیون‌ها virtual thread این آنتی‌پترن است: هر نخ یک نسخه یعنی مصرف حافظهٔ انفجاری. با virtual threads یا شیء را per-task بساز، یا آن را pool کن، یا ScopedValue را در نظر بگیر.

به‌روزرسانی مدرن — pinning تقریباً حل شد

در فصل گفتیم synchronized باعث pinning می‌شود. JEP 491 در JDK 24 دقیقاً این را رفع کرد: بلاک synchronized دیگر معمولاً carrier را گروگان نمی‌گیرد. فقط native/JNI و چند گوشهٔ نادر مانده‌اند. برای پایش pinning، به‌جای فلگ قدیمیِ -Djdk.tracePinnedThreads (که در JDK 24 deprecate شد) از رویداد JFR به‌نام jdk.VirtualThreadPinned استفاده کن. تعداد carrierها را هم با -Djdk.virtualThreadScheduler.parallelism تنظیم می‌کنی.

ساخت دستی با Thread.ofVirtual().name("job-", 0).start(runnable) یا Thread.startVirtualThread(r). دو همراهِ مدرن الگوی fan-out را کامل می‌کنند: structured concurrency (در JDK 25 هنوز preview، JEP 505، حالا با StructuredTaskScope.open() و Joiner) چند subtask را در یک scope اجرا می‌کند تا همه با هم fail/cancel شوند و leak نخ نداشته باشی؛ و ScopedValue (در JDK 25 نهایی شد، JEP 506) جایگزینِ immutable و ارزان‌ترِ ThreadLocal برای انتقال کانتکست به subtaskهاست — بدون نیاز به remove()، چون دامنه‌اش خودکار بسته می‌شود.

سؤال‌های سختِ سنیور

۱. [سخت] taskی روی نخ دیگر مقداری را در یک فیلد معمولی (بدون volatile) می‌نویسد و تو بعد از `future.get()` می‌خوانی. آیا به volatile نیاز داری؟

نه. JMM تضمین می‌کند هرچه task انجام می‌دهد happens-before بازگشت Future.get() است، و هرچه قبل از submit نوشتی happens-before شروع task. پس مرز submit→run→get خودش visibility کامل می‌سازد و نیازی به volatile/synchronized برای انتقال نتیجه نیست. اما اگر بدون هیچ نقطهٔ همگام‌سازی (نه get، نه queue، نه latch) از fire-and-forget نتیجه بخوانی، هیچ ضمانتی نداری و مقدار کهنه محتمل است.

۲. [سخت] pod تو در k8s با `cpu limit: 1` روی نودی ۳۲core اجرا می‌شود. `newFixedThreadPool(Runtime.getRuntime().availableProcessors()*2)` و parallel streamها رفتار عجیب دارند. چرا؟

چون JVM از JDK 10+ container-aware است و cgroup را می‌خواند: با limit یک core، availableProcessors() مقدار ۱ (نه ۳۲) برمی‌گرداند. پس pool تو ۲ نخ می‌شود، و مهم‌تر، commonPool() که parallelism‌اش cores-1 است عملاً یک‌نخی می‌شود و parallel streamها سریال اجرا می‌شوند. راه‌حل: عدد واقعی را لاگ کن، در صورت نیاز -XX:ActiveProcessorCount را صریح ست کن، و cpu request == limit بگذار.

۳. [سخت] یک endpoint با `parallelStream().map(remoteCall)` باعث کندی درخواست‌های کاملاً بی‌ربط می‌شود. چرا و دو راه‌حل؟

چون parallel stream روی ForkJoinPool.commonPool() سراسری اجرا می‌شود و فراخوانی blockingِ remote، نخ‌های این pool را قفل می‌کند؛ چون این pool در کل JVM مشترک است، بقیهٔ استفاده‌کننده‌ها گرسنه می‌مانند. راه‌حل ۱: stream را داخل یک ForkJoinPool اختصاصی submit کن. راه‌حل ۲ (بهتر): از parallel stream برای کار blocking استفاده نکن؛ به‌جایش CompletableFuture.supplyAsync(task, dedicatedIoPool) یا virtual threads. اگر مجبوری داخل FJP block کنی، ManagedBlocker بزن تا pool worker جبرانی بسازد.

۴. [سخت] poolِ کراندار با ۱۰ نخ داری؛ هر task خودش subtask به همان pool submit و روی `get` block می‌کند. زیر بار برنامه hang می‌کند. تشخیص؟

Thread-starvation deadlock. اگر ۱۰ taskِ والد همزمان هر ۱۰ نخ را بگیرند و هرکدام منتظر subtaskِ خودش باشد، هیچ نخی برای اجرای subtaskها نمی‌ماند؛ والدها منتظر فرزندان، فرزندان منتظر نخ آزاد. راه‌حل: poolهای جدا برای لایه‌های مختلف، یا زنجیرهٔ non-blocking با thenCompose، یا virtual threads که نخ در آن‌ها کمیاب نیست.

۵. `thenApply` در برابر `thenApplyAsync` — هرکدام callback را کجا اجرا می‌کند و کجا می‌زندت؟ فرق `get` و `join`؟

thenApply: callback روی همان نخی که stage قبلی را complete کرده (یا حتی نخ فراخواننده اگر از قبل complete بوده) — برای کار سنگین غیرقابل‌پیش‌بینی. thenApplyAsync بدون executor: روی commonPool (دام قحطی). thenApplyAsync با executor: روی pool تو — حالت درست prod. get() استثنای checked می‌دهد (ExecutionExceptionjoin() نسخهٔ unchecked است (CompletionException) و داخل زنجیره تمیزتر.

۶. [سخت] یک متد `@Async` را از داخل همان کلاس صدا می‌زنی ولی سنکرون اجرا می‌شود. چرا؟ و چرا خطاهای `@Async void` گم می‌شوند؟

چون @Async با proxy کار می‌کند؛ فراخوانی داخلی (self-invocation) از proxy رد نمی‌شود پس async‌سازی اتفاق نمی‌افتد و کد سنکرون می‌شود — باید از bean دیگری صدایش بزنی یا proxy را self-inject کنی. دربارهٔ خطا: متد @Async با بازگشت void مثل execute بدون Future است؛ استثنا هیچ‌جا catch نمی‌شود و در لاگ گم می‌شود. راه‌حل: AsyncUncaughtExceptionHandler تعریف کن یا بازگشت را CompletableFuture کن تا خطا در join/get دربیاید.

۷. [سخت] با virtual threads دیگر pool را sizing نمی‌کنی؛ پس چطور جلوی downstreamی که فقط ۲۰ اتصال همزمان می‌دهد را می‌گیری؟ و pinning را حالا چطور تشخیص می‌دهی؟

concurrency را با یک Semaphore(20) دورِ فراخوانی downstream bound می‌کنی؛ Semaphore در دنیای virtual thread همان نقش maximumPoolSize را دارد. برای pinning: در JDK 24 با JEP 491 دیگر synchronized معمولاً pin نمی‌کند؛ فقط native/JNI مانده. برای پایش از رویداد JFR به‌نام jdk.VirtualThreadPinned استفاده کن (فلگ قدیمی -Djdk.tracePinnedThreads در JDK 24 deprecate شد).

جمع‌بندی سنیور
  • مرز submit → run → Future.get خودش happens-before می‌سازد؛ برای انتقال نتیجه به volatile نیازی نیست، ولی fire-and-forget بدون نقطهٔ همگام‌سازی امن نیست.
  • در container، availableProcessors() از cgroup می‌آید و ممکن است بسیار کوچک‌تر از coreهای node باشد — همهٔ sizingها را می‌شکند.
  • commonPool() و parallel streamها سراسری‌اند؛ blocking داخلشان کل JVM را قحطی‌زده می‌کند (executor اختصاصی یا ManagedBlocker). pool تودرتوی کراندار هم thread-starvation deadlock است.
  • در CompletableFuture نسخهٔ Async با executor خودت را بده؛ pool را observable کن و MDC را با decorator منتقل کن.
  • دام‌های @Async: self-invocation سنکرون می‌شود، void خطا را می‌بلعد، default ممکن است pool نداشته باشد.
  • با virtual threads: Semaphore جای sizing، ThreadLocalِ گران ممنوع، pinning در JDK 24 تقریباً حل شد؛ ScopedValue و structured concurrency الگوی مدرن fan-out‌اند.
جمع‌بندی
  • نخِ سکویی گران است (پشتهٔ ~۱MB + فراخوان سیستمی)؛ برای همین Executorها ثبتِ وظیفه را از ساختِ نخ جدا می‌کنند.
  • شش حالتِ Thread.State را حفظ کن. BLOCKED فقط رقابت بر سرِ مانیتور است؛ I/O در سطح OS به‌صورت RUNNABLE دیده می‌شود.
  • run() نخ نمی‌سازد، start() می‌سازد. Runnable نتیجه ندارد، Callable دارد و نتیجه‌اش را با Future.get() می‌گیری — که مسدودکننده است و استثنا را تا لحظهٔ get() قایم می‌کند.
  • در تولید ThreadPoolExecutor را دستی با صفِ کران‌دار و سیاستِ ردّ بساز. با صفِ بی‌کران، maximumPoolSize بی‌اثر می‌شود.
  • CPU-محور را به تعداد هسته اندازه بگیر؛ I/O-محور را با cores × (1 + wait/compute)، و در گلوگاهِ واقعی سقف بگذار.
  • برای زمان‌بندی، فرقِ fixed-rate و fixed-delay را بدان و بدنهٔ وظیفه را در try/catch بپیچ (وگرنه مرگِ خاموش).
  • همیشه shutdown()awaitTermination()shutdownNow(). لغو تعاونی است؛ InterruptedException را هرگز نبلع.
  • در استخرها هر ThreadLocal را در finally با remove() پاک کن (درستی + امنیت + حافظه).
  • در جاوا ۲۱+ برای کارِ I/O-محور نخ‌های مجازی به‌ازای هر وظیفه بساز و اندازه‌گذاری را فراموش کن — فقط مراقبِ pinning باش و آن‌ها را استخر نکن.

منابع: ThreadPoolExecutor (JDK 17)، Managing Throughput with Virtual Threads — inside.java، RejectedExecutionHandler — Baeldung

Come sit next to me. We are about to talk about one of those topics everyone dreads in interviews and bleeds over in production: concurrency. But I promise you — if you truly understand three things, this topic stops being scary magic and becomes something you hold in the palm of your hand. This whole chapter is those three things, built from the ground up, one small brick at a time.

First, a mental picture: a thread is like a worker who can follow one task. One worker means tasks happen one after another. Several workers means tasks can go forward at the same time. The entire argument of this chapter is about how you build those workers, how you hand them work, and how you politely say "that's enough, go home."

Roadmap

Here is the path we'll walk:

  1. What a thread actually is and why creating one is expensive (that expense is the whole reason Executors exist).
  2. The thread lifecycle — the six states a thread can be in.
  3. The unit of work: Runnable, Callable, and the result-handle Future.
  4. The Executor framework and its heart, ThreadPoolExecutor, with its seven parameters.
  5. Pool-sizing math — where senior candidates earn their stripes.
  6. Scheduling (ScheduledExecutorService) and graceful shutdown.
  7. Cooperative cancellation (InterruptedException) and ThreadLocal, which leaks in pools.
  8. Daemon threads, and finally virtual threads (Java 21+), which make the sizing math evaporate. At the end: fifteen interview questions with full answers. Ready? Let's go.

Part 0 — words you must know first

Before we jump into code, let's plant a few words we'll keep reusing, each with a small analogy. Keep these in the back of your mind and the rest of the chapter is easy.

A restaurant and its cooks

Picture a restaurant. A cook is a thread: one person who can carry one order from start to finish. The maître d' who decides which cook works right now is the OS scheduler — because the kitchen only has a few stoves (CPUs/cores), not everyone can cook simultaneously, so the scheduler takes turns. Each cook's personal station with its board, knife, and spices is the stack: memory that belongs only to that thread. Walking to the storeroom to fetch a new cook is a system call: slow and costly. That's why serious restaurants don't hire and fire cooks mid-shift; they keep a fixed team and spread orders across them. That's exactly what an Executor does for you.

  • Thread: an independent unit of execution. A thread drives one sequence of instructions.
  • OS scheduler: the part of the operating system that decides which thread runs on which core and for how long. You don't control it directly; you only make requests.
  • Stack: each thread's private memory for local variables and the chain of method calls. Every thread has its own (typically ~1 MB reserved).
  • System call: when your program asks the OS kernel for something (e.g., "make a new thread"). These are relatively slow.
  • Blocking: when a thread stops and waits for something to happen — e.g., waiting for a database reply. While blocked, that thread does nothing but wait.
  • Monitor lock: a lock acquired on an object via the synchronized keyword so only one thread at a time enters a critical section. Like a bathroom key only one person holds.
  • CPU-bound vs I/O-bound: CPU-bound work means the thread is truly computing (compression, hashing). I/O-bound work means the thread mostly waits on something external (DB, network, disk). This distinction is the backbone of pool sizing.

Now that we have the tools, let's meet the thread itself.

Mental model: what a thread is and why it's expensive

A java.lang.Thread in Java is really a thin wrapper over an OS thread — we call that OS-level thread a platform thread, to distinguish it later from the virtual thread.

Why do I say it's expensive? When you create a thread, three costly things happen:

  1. A large native stack is allocated (typically ~1 MB reserved).
  2. The thread is registered with the OS scheduler.
  3. All of this requires a system call.
Why do we even have Executors?

That very expense of creating a thread is the entire reason the Executor framework exists. Instead of hiring a fresh cook for every order (slow and costly), you keep a small fixed team of cooks and spread work across them. This is called decoupling task submission from thread creation. You hand off the work; which ready thread runs it is no longer your concern.

Three questions frame this whole chapter. Answer these three and concurrency stops being mysterious:

  1. What is the unit of work? Runnable (no result, no checked exception), Callable<V> (returns a value, may throw a checked exception), or the raw Thread subclass (almost never the right answer).
  2. Who runs it, and on what thread? A bare Thread, or an ExecutorService backed by a pool.
  3. How does it stop? Cooperatively, via interruption — because Java has no safe preemptive Thread.stop().

Let's unpack these three one by one.

Thread lifecycle & states

Over its life, a thread moves between several "states." Java keeps these in an enum called Thread.State. This enum is the ground truth — memorize it, because interviewers almost always probe the difference between WAITING and BLOCKED.

An employee through the day

Picture an employee. NEW: just hired, hasn't started working yet. RUNNABLE: sitting at the desk, working or ready to work. BLOCKED: standing outside the conference room, waiting for the previous person to come out and hand over the key (the lock). WAITING: waiting for a colleague to signal them, with no deadline. TIMED_WAITING: same waiting, but with an alarm clock — "I'll wait until 3 o'clock." TERMINATED: done for the day and gone home.

State Meaning How you get there
NEW Created, not yet started new Thread(...) before start()
RUNNABLE Eligible to run or running (JVM does not distinguish "ready" vs "on-CPU") after start()
BLOCKED Waiting to acquire a monitor lock (synchronized) contending for a monitor
WAITING Waiting indefinitely for another thread's signal wait(), join(), LockSupport.park()
TIMED_WAITING Same, but with a deadline sleep(ms), wait(ms), join(ms), park(ns)
TERMINATED run() returned or threw task finished

Now two traps, exactly where the interviewer puts a finger:

A thread waiting on I/O is RUNNABLE, not BLOCKED

This one surprises many people. A thread blocked on an I/O operation (say, stuck mid-read on a network socket, waiting for data) shows up in a thread dump as RUNNABLE, not BLOCKED. Why? Because that block happens at the OS level and the JVM cannot see it at all — from the JVM's point of view the thread is "able to run." Java uses the BLOCKED label only and specifically for contention over a synchronized monitor lock. So next time you see a thread dump full of RUNNABLE, don't be fooled; those threads may actually be waiting on the network.

The classic trap: run() instead of start()

start() is one-shot. If you call start() twice you get an IllegalThreadStateException. And crucially: if you call run() directly, no new thread starts at all — the body executes on the current thread, synchronously (right here, right now). Meaning no concurrency happens whatsoever. This is one of the most common accidental bugs beginners hit.

Runnable task = () -> System.out.println("on " + Thread.currentThread());
Thread t = new Thread(task);
t.run();    // BUG: runs on the *current* thread, no concurrency
t.start();  // correct: spawns a new thread

Keep the difference like this: run() is just an ordinary method that you call; it's start() that tells the OS "create a fresh thread and let that thread call run()."

Thread vs Runnable vs Callable/Future

Now the first question: what is the unit of work?

A golden rule: prefer composition over inheritance. That is, instead of extending Thread (extends Thread), write your work as a Runnable or Callable and hand it to a thread. Why? Because subclassing Thread couples your work to a thread forever and prevents you from reusing it in a pool. But a Runnable is a self-contained package of work that any thread can pick up and run.

An order ticket vs a dedicated cook

Runnable/Callable is like an order ticket: on paper you've written "make one pizza." That ticket can be handed to any cook. But extends Thread is like building, for each order, a cook specific to that order with the work sewn into their body — useless for the next order. The order ticket is reusable; the sewn-in cook is not.

The difference between Runnable and Callable:

  • Runnable means "fire and forget": returns no value and its signature allows no checked exceptions.
  • Callable<V> means "do something and give me the result back": returns a value of type V and may throw a checked exception.
// Runnable: fire-and-forget, no return, no checked exceptions in signature
Runnable r = () -> log.info("done");

// Callable: returns a value, may throw checked exceptions
Callable<Integer> c = () -> {
    return httpClient.fetchCount(); // may throw IOException
};

So if Callable returns a result but doesn't run right now — it finishes later on another thread — how do I get the result? That's where Future enters.

A dry-cleaning ticket

Future<V> is like the ticket the dry cleaner hands you. You've dropped off the work (the clothes), the result isn't ready yet, but you hold a handle you can use to collect the result later. When you call f.get(), it's like bringing the ticket back and saying "give me my clothes now" — and if they're not ready yet, you stand there and wait (that's what "get() is blocking" means).

ExecutorService pool = Executors.newFixedThreadPool(4);
Future<Integer> f = pool.submit(c);
try {
    Integer n = f.get(2, TimeUnit.SECONDS); // blocks up to 2s
} catch (ExecutionException e) {
    // the task threw; e.getCause() is the ORIGINAL exception
    Throwable real = e.getCause();
} catch (TimeoutException e) {
    f.cancel(true); // true => interrupt the running thread
}

Three things about Future that trip people up — read these carefully, because all three can create silent bugs in production:

Exceptions hide until get()

This is Future's most dangerous behavior. If a task you handed off with submit() throws mid-work, nothing visibly happens. No log, no error, nothing. The exception is wrapped in an ExecutionException and delivered to you only when you call get() (with the original exception in e.getCause()). By contrast, if you hand off work with execute(Runnable) — which gives no Future at all — an uncaught exception goes to that thread's UncaughtExceptionHandler and usually gets printed. This asymmetry silently buries bugs. Rule: either always get() your Futures, or deliberately use execute.

  • f.cancel(true) only sets the interrupt flag; the task itself must cooperate (I'll explain in the interruption section below). cancel(false) prevents a not-yet-started task from running but won't interrupt a running one.
  • get() is blocking. The calling thread stands there until the result is ready. If you want to chain work without standing still, prefer CompletableFuture, which is non-blocking and composable (with methods like thenApply, thenCompose, allOf).

The Executor framework

Now the second question: who runs the work? The professional answer: an Executor.

The Executor framework has a clean hierarchy, going from simple to complete:

  • Executor — the most basic, has just one method: execute(Runnable).
  • ExecutorService — extends it, adding submit, invokeAll, and lifecycle management (shutdown).
  • ScheduledExecutorService — extends that further, adding delays and periodic execution.

But the real workhorse behind all of these is the class ThreadPoolExecutor. First, let's see why you shouldn't blindly trust the ready-made shortcuts.

The Executors factory — and why to distrust it

The Executors class has several convenient factory methods that build a pool for you in one line. The catch: several are dangerous in production, because they use either unbounded queues or unbounded thread counts:

Executors.newFixedThreadPool(n);   // LinkedBlockingQueue with NO capacity bound → OOM risk
Executors.newSingleThreadExecutor();// same unbounded queue
Executors.newCachedThreadPool();    // maxPoolSize = Integer.MAX_VALUE → unbounded threads

Let me make the danger concrete:

Why ready-made factories blow up in production

"Unbounded" means "no ceiling." Under sustained heavy load, newFixedThreadPool grows its queue of waiting work without limit — work piles on work until the heap fills and you get an OutOfMemoryError. On the other side, newCachedThreadPool has no cap on thread count (Integer.MAX_VALUE), so under pressure it spawns threads until the OS refuses to make more and the app falls over. The senior practice is one sentence: construct ThreadPoolExecutor directly and by hand, with a bounded queue and an explicit rejection policy. That way, when the system is under pressure, instead of exploding it politely says "no."

ThreadPoolExecutor: the seven parameters

Let's look at the full constructor. It takes seven parameters, and each is an engineering decision:

ThreadPoolExecutor pool = new ThreadPoolExecutor(
    4,                              // corePoolSize
    16,                             // maximumPoolSize
    60L, TimeUnit.SECONDS,          // keepAliveTime for threads above core
    new ArrayBlockingQueue<>(1000), // bounded work queue
    new ThreadFactoryBuilder()      // name threads! (Guava, or a hand-rolled factory)
        .setNameFormat("api-worker-%d").build(),
    new ThreadPoolExecutor.CallerRunsPolicy() // rejection handler
);
  • corePoolSize (4): the number of always-kept threads, retained even when idle.
  • maximumPoolSize (16): the most threads the pool is allowed to grow to.
  • keepAliveTime (60s): threads above core that sit idle this long are torn down.
  • work queue (ArrayBlockingQueue capacity 1000): where tasks wait for a free thread. Making this queue bounded is critical — you'll see why in a moment.
  • ThreadFactory: the factory that creates threads; here we use it to name them.
  • RejectedExecutionHandler: what to do with a new task when everything is full.
The submission algorithm, like a café

Imagine you run a café. corePoolSize is your permanent baristas (say 4). The queue is the waiting chairs (1000 of them). maximumPoolSize is the most baristas you can call up from the back (16). Now when a customer arrives: first, if you have fewer than your permanent baristas busy, one steps up. Then, if all are busy, the customer sits in a waiting chair. Only when all the chairs are full too do you call up an extra barista from the back. And if baristas have hit the max and the chairs are full, you tell the customer "no" at the door.

The submission algorithm of ThreadPoolExecutor is exactly this order and is worth memorizing, because it explains a notorious surprise:

  1. If running threads < corePoolSize, start a new thread for the task (even if other threads are idle).
  2. Else, try to enqueue the task.
  3. If the queue is full, and running threads < maximumPoolSize, start a new (non-core) thread.
  4. If the queue is full and threads == max, reject via the handler.
The big surprise: with an unbounded queue, maximumPoolSize is ignored!

Here's the trap. If your queue is unbounded (like newFixedThreadPool), step 2 always succeeds — the queue never fills. So the "queue is full" condition in step 3 never becomes true, which means step 3 never fires and maximumPoolSize is completely dead. The pool never grows past corePoolSize, no matter how heavy the load. This surprises almost everyone. So if you want your pool to scale up under load, you must give it a bounded queue (or a SynchronousQueue, which has zero capacity — so every task either immediately demands a new thread or is rejected; that's exactly how newCachedThreadPool works).

Rejection policies (RejectedExecutionHandler)

When the pool is fully saturated (queue full and threads at max), what do we do with a new task? You have four ready-made policies:

Policy Behavior Use when
AbortPolicy (default) Throws RejectedExecutionException You want fail-fast and a caller that handles overflow
CallerRunsPolicy Runs the task on the submitting thread Backpressure: the submitter slows down because it's busy running the task, naturally throttling intake
DiscardPolicy Silently drops the task Only when losing work is acceptable (e.g., best-effort metrics)
DiscardOldestPolicy Drops the oldest queued task, retries submit Rarely correct; drops work that's been waiting longest
What is backpressure?

"Backpressure" means that when the system is under strain, instead of blindly accepting more and more work until it bursts, it pushes back on the source: "send slower." CallerRunsPolicy gives you this for free: when the pool is full, the very thread that submitted the task is forced to run it. As long as it's busy running that task, it cannot submit new work — so the intake rate drops by itself. This is the pragmatic default for request pipelines. One warning though: if that submitting thread is the same thread accepting HTTP connections, running a task on it means accepting new connections stops for a while. Which may or may not be what you want.

Pool-sizing math

Now we reach where senior candidates truly earn their stripes: how many threads should I use? The answer depends on one thing — is your work CPU-bound or I/O-bound?

CPU-bound case: CPU-bound work (compression, hashing, pure computation) keeps a core busy the whole time. Here, if you make more threads than cores, you gain nothing; you only add context-switch overhead (the cost the OS pays to set one thread aside and put another on the core):

threads ≈ number_of_cores        (or cores + 1 to cover the occasional page fault)
Why more threads than cores doesn't help

If you have 4 stoves and 4 cooks, one cook per stove, everything's perfect. Now bring 8 cooks to the same 4 stoves. Do you cook faster? No! There are still only 4 stoves. The extra cooks either stand idle or constantly swap places at the stove, and that swapping itself takes time. CPU-bound work means "the stove is always lit"; so peak speed is when the number of cooks equals the number of stoves.

I/O-bound case: I/O-bound work spends most of its time waiting (on the DB, HTTP, disk). Here the story is completely different: you want enough threads that while several are waiting, others can use the CPU. Brian Goetz's famous formula (from Java Concurrency in Practice), which is really a restatement of Little's Law:

threads = cores × targetUtilization × (1 + waitTime / computeTime)

Let's pin it with a worked example. Suppose 8 cores, target 100% utilization, and each task waits 90 ms on the DB and computes for 10 ms (so wait/compute = 9):

threads = 8 × 1.0 × (1 + 90/10) = 8 × 10 = 80 threads
The wait/compute ratio is the single most important sizing number

See what happened? An I/O-heavy service legitimately wants about 80 threads on 8 cores — ten times the number of cores! The logic is simple: since each thread spends 90% of its time just waiting, you need many threads so that the cores never sit idle and one of them is always doing useful work. Undersize it and the CPU sits idle while requests queue; oversize it and you burn memory (~1 MB stack each) and thrash the system (constant, useless swapping). This wait/compute ratio is the single most important number in pool sizing — and it's exactly the pressure that virtual threads (at the end of the chapter) relieve.

int cores = Runtime.getRuntime().availableProcessors();
double waitToCompute = 9.0; // measured from profiling, not guessed
int size = (int) (cores * (1 + waitToCompute));

One important practical note: you must get the waitToCompute number from real profiling, not from a guess.

Always cap at the real bottleneck

The formula gives you a number, but the real world has another hard ceiling. If your DB connection pool has only 20 connections, more than ~20 concurrent DB threads is useless — thread number 21 just queues behind the connection pool. So always bound your pool to the real downstream bottleneck. Measure where it actually gets narrow, and stop there.

ScheduledExecutorService

Sometimes you don't want the work to run right now; you want it "in 5 seconds" or "every 10 seconds." For these you have ScheduledExecutorService.

Why Timer is effectively dead

You may know the old Timer class. It's considered obsolete in practice, for two reasons: (1) it uses just one thread, so if one task is slow the rest fall behind; and (2) if a task throws an exception, the whole Timer silently dies. ScheduledExecutorService is its proper, resilient replacement.

ScheduledExecutorService sched = Executors.newScheduledThreadPool(2);

sched.schedule(() -> log.info("once, after 5s"), 5, TimeUnit.SECONDS);

// fixed RATE: fires every 10s regardless of task duration (can pile up / overlap-guarded)
sched.scheduleAtFixedRate(this::poll, 0, 10, TimeUnit.SECONDS);

// fixed DELAY: waits 10s AFTER each completion — no pile-up
sched.scheduleWithFixedDelay(this::poll, 0, 10, TimeUnit.SECONDS);

You should feel the difference between these two methods in your bones, because it's a critical distinction:

A clockwork train vs a spaced-out bus

scheduleAtFixedRate is like a train whose departure time is fixed: it leaves on the hour every hour, no matter how long the previous trip took. If one run takes longer than the 10-second period, the next run starts immediately after it to catch up (but they never run concurrently on the same task). scheduleWithFixedDelay, by contrast, is like a bus that rests for 10 seconds after returning to the depot each time, then sets off again — so the gap is always between the end of one and the start of the next, and runs never pile up.

  • scheduleAtFixedRate: targets a fixed period. If a run exceeds the period, subsequent runs execute back-to-back to catch up.
  • scheduleWithFixedDelay: inserts a fixed gap between the end of one run and the start of the next, so runs never overlap.
The silent-death trap

This one can genuinely eat hours of your life. If a scheduled task throws an uncaught exception, that task is cancelled and never runs again — the pool silently suppresses it and no further executions happen, with no error in the logs. Picture a cron that ran for months and one day "just stops." This is almost always the cause. Always wrap the body of periodic tasks in try/catch:

sched.scheduleWithFixedDelay(() -> {
    try { doWork(); }
    catch (Throwable t) { log.error("poll failed, will retry next tick", t); }
}, 0, 10, TimeUnit.SECONDS);

Shutdown: shutdown() vs shutdownNow()

An important note: you must shut down an ExecutorService yourself, or your JVM won't exit at all! Because pool worker threads are by default non-daemon (user) threads, and while alive they keep the JVM alive. (I'll explain daemon below.)

You have two ways to shut down:

Method Accepts new tasks? Already-queued tasks? Running tasks?
shutdown() No Run to completion Finish
shutdownNow() No Returned as List<Runnable>, not run Interrupted
Politely closing vs immediately closing the restaurant

shutdown() means "we lock the front door, take no new customers, but everyone already inside who ordered gets their meal in full." shutdownNow() means "closed right now! Cooks, drop what you're doing (interrupt), and orders not yet started stay on the counter" — it hands those unstarted orders back to you as a list so you can decide what to do with them.

Note that shutdownNow() interrupts running tasks (sets their interrupt flag) but cannot force-stop a task that ignores interruption. The canonical, correct graceful-shutdown idiom is this:

void shutdownGracefully(ExecutorService pool) {
    pool.shutdown(); // stop accepting, let queued tasks finish
    try {
        if (!pool.awaitTermination(30, TimeUnit.SECONDS)) {
            pool.shutdownNow(); // force: interrupt running tasks
            if (!pool.awaitTermination(10, TimeUnit.SECONDS))
                log.error("pool did not terminate");
        }
    } catch (InterruptedException e) {
        pool.shutdownNow();
        Thread.currentThread().interrupt(); // restore the flag
    }
}
awaitTermination doesn't shut down + Java 19 and try-with-resources

Two subtle points: first, awaitTermination returns only a boolean and does not itself trigger shutdown — you must call shutdown() or shutdownNow() first, otherwise you'll wait forever. Second, since Java 19, ExecutorService implements AutoCloseable; its close() method behaves like "shutdown() then await," so you can put the pool in a try-with-resources and rest easy that it closes automatically.

The interruption model (cooperative cancellation)

We've reached the third question: how does a thread stop? And here Java made a deliberate, important decision: there is no safe way to forcibly kill a thread.

Why Thread.stop() was deprecated

You might ask, "don't we have a Thread.stop() method?" We do, but it's deprecated and you must never use it. The reason is dangerous: stop() cuts the thread off mid-work and releases its locks mid-mutation — that is, in the middle of a half-finished change to an object. The result? Objects left in a corrupt, inconsistent state. Like yanking a surgeon out of the room mid-operation. That's why Java closed this door.

So cancellation in Java is cooperative. You only request cancellation (via interrupt()), and the target thread itself chooses to notice and cleanly wind down.

Tapping someone on the shoulder

interrupt() is like gently tapping a colleague who's wearing headphones on the shoulder and saying "whenever you can, finish up and come over." You can't switch off their brain; you only send a signal. If they're cooperative, they notice at the first opportunity and wind down cleanly. If they're oblivious and never check, nothing happens. Cooperative cancellation is exactly this: it needs the target thread's cooperation.

Interruption is really just a per-thread boolean flag, plus a few rules:

  • thread.interrupt() — sets the flag (sends the signal).
  • Thread.interrupted()static; reads and clears the flag on the current thread.
  • thread.isInterrupted() — reads the flag without clearing it.
  • Blocking methods that declare throws InterruptedException in their signature (sleep, wait, join, BlockingQueue.take, Future.get) clear the flag and throw an InterruptedException when interrupted.

Now, when you catch an InterruptedException, you have exactly two correct responses:

  1. Propagate it (let it bubble up to higher layers), or
  2. Restore the flag and exit the work, so callers up the stack can still see the cancellation:
try {
    queue.take();
} catch (InterruptedException e) {
    Thread.currentThread().interrupt(); // NEVER swallow silently
    return; // stop the unit of work
}
The cardinal sin: swallowing InterruptedException

The worst thing you can do is catch InterruptedException with an empty catch and do nothing. Remember: when this exception is thrown, the interrupt flag has already been cleared. Now if you swallow it too, you've completely destroyed the one signal that told the world "stop" — and now your task can no longer be cancelled, and even shutdownNow can't stop it. This is one of the most-flagged bugs in senior code review. Never, ever swallow it silently.

One last note: blocking methods throw InterruptedException themselves, but if you have a CPU-bound loop that calls no blocking method, you must check the flag manually:

while (!Thread.currentThread().isInterrupted()) {
    doOneChunk();
}
// loop exits cleanly on interrupt — cooperative cancellation

ThreadLocal — and why it leaks in pools

ThreadLocal<T> gives each thread its own private copy of a variable. It's the standard mechanism for holding "per-request context" — like a transaction ID, the current user, or a SimpleDateFormat instance — without having to pass it by hand through every method.

Each employee's personal locker

ThreadLocal is like each employee having a personal locker at work. Everyone has access to "the locker" (the variable is shared), but the inside of each person's locker is theirs alone and no one sees another's contents. This is great, until… the employees don't change. If you later give that same locker to a new employee without emptying it, the previous person's belongings are still inside. And that is exactly what happens in a thread pool.

private static final ThreadLocal<SimpleDateFormat> FMT =
    ThreadLocal.withInitial(() -> new SimpleDateFormat("yyyy-MM-dd"));

ThreadLocal values are stored in a map inside the Thread object itself, keyed by the ThreadLocal object. In pooled threads this creates two serious hazards:

Hazard 1: value leakage and stale data (correctness and security bug)

Pool threads are long-lived and reused. A value you set while serving request A is still there when the same thread later serves request B. Now request B sees request A's context — for example, user A's identity! This is both a correctness bug and a serious security bug. The iron rule: if you set a ThreadLocal in a pooled thread, you must remove() it in a finally block:

try {
    RequestContext.CURRENT.set(ctx);
    handle(request);
} finally {
    RequestContext.CURRENT.remove(); // mandatory in pools
}
Hazard 2: memory leak

ThreadLocal's internal map (the ThreadLocalMap) uses weak references for its keys (the ThreadLocal object itself) but strong references for the values. Now think: if the ThreadLocal object is no longer referenced anywhere but the thread stays alive (which is exactly how pool threads behave), the value can't be collected by the garbage collector — until that map slot is reused or the thread dies. If you've cached large objects in a ThreadLocal and your pool is long-lived, the heap slowly fills. remove() fixes this too.

InheritableThreadLocal doesn't help you in pools

You might think InheritableThreadLocal is the answer — this variant copies the value to a child thread at creation time. But pool threads aren't re-created per task (that's the whole point of a pool!), so it does not propagate context into pool tasks. To propagate context correctly in pools you need explicit wrapping, or ScopedValue, which arrived in Java 21.

Daemon threads

Every thread in Java is either a user thread or a daemon thread. The key rule: the JVM exits precisely when the last non-daemon thread finishes — and at that moment, it abruptly abandons any daemon thread still running (with no guarantee that finally blocks run).

A permanent employee vs a night watchman

A user thread is like a permanent employee: as long as one has work, the building (the JVM) stays open. A daemon thread is like a night watchman or the HVAC system: it does background work, but the moment the last permanent employee leaves, the lights go out and the watchman too — like it or not, even mid-task — is left outside. That's why you never assign critical work to the night watchman.

Thread t = new Thread(task);
t.setDaemon(true);   // must be set BEFORE start(), else IllegalThreadStateException
t.start();
  • Use daemon threads for background infrastructure that should never keep the app alive: metrics flushers, cache-eviction timers, heartbeat pollers.
  • Never use a daemon thread for work that must complete or must clean up (writing a file, flushing a transaction) — the JVM may kill it mid-write.
  • A thread inherits its parent's daemon status by default. That's why Executors threads are user (non-daemon) threads and keep the JVM alive until you shut the pool down. If you want daemon threads in a pool, you must supply a ThreadFactory that calls setDaemon(true).

Virtual threads (Java 21+) — where the sizing math evaporates

And now the biggest change of recent years. Java 21 made virtual threads permanent. Let me build the intuition first.

Motorbike couriers and company cars

Imagine you run a courier company. A platform thread is like an expensive company car: there are only a few, and each is costly. If a courier reaches a door and has to wait for it to open, that expensive car sits parked and idle right there. A virtual thread is like a cheap courier: when they reach a door and must wait, they get off the car (unmount) and free it for another courier to hop on; when the door opens, they mount a free car again. There are few cars (the platform carrier threads = carriers), but there can be millions of couriers.

So a virtual thread is a thread that the JVM itself schedules onto a small pool of carrier platform threads; when it blocks on I/O, instead of blocking an OS thread it unmounts — it parks in user space and frees its carrier. They're cheap enough that you can create millions.

// Each task gets its own virtual thread; NOT a pool to be sized
try (var exec = Executors.newVirtualThreadPerTaskExecutor()) {
    for (var req : requests)
        exec.submit(() -> handleBlocking(req)); // blocking calls are fine here
} // close() awaits all tasks (AutoCloseable)
The paradigm shift: you no longer size a pool

This is the heart of it. For I/O-bound work, you no longer size a pool. You write straightforward blocking code, one virtual thread per task, and let the JVM multiplex them onto the few carriers. Remember Goetz's formula? That formula was really a workaround for the "expensive platform thread" problem. Virtual threads remove, at the root, the very constraint the formula was compensating for.

But a few things still bite:

Pinning — the silent killer of virtual threads

If a virtual thread is inside a synchronized block (or in the middle of a native/JNI call) and blocks there, it can no longer unmount from its carrier — it gets pinned to that carrier, i.e., nailed to it. If enough threads get pinned to the carriers, the carrier pool starves and throughput collapses. The fix: around blocking I/O, use a ReentrantLock instead of synchronized, which lets the virtual thread unmount. (Later JDKs, 24+, remove most synchronized-induced pinning, but on Java 21 itself the problem is real.)

  • Don't pool them. Making a newFixedThreadPool of virtual threads defeats the whole philosophy; always create one per task.
  • They give no benefit for CPU-bound work. A virtual thread helps with waiting, not computing. For pure computation, keep the platform-thread pool sized to your cores.

Common pitfalls & gotchas (rapid fire)

Let's review all the traps we saw in this chapter in one place — these are exactly the ones that cause pain in production:

  • Calling run() instead of start() — no new thread.
  • Unbounded newFixedThreadPool queue → OOM under load.
  • maximumPoolSize ignored because the queue is unbounded.
  • Swallowing InterruptedException → uncancellable tasks.
  • Forgetting ThreadLocal.remove() in a pool → leaked/stale context.
  • Not naming threads → useless thread dumps in production. Always supply a ThreadFactory.
  • Exceptions in submitted tasks silently lost until get().
  • Scheduled task throws once → silently stops forever.
  • Sharing a non-thread-safe object (SimpleDateFormat, Random under contention) across pool threads.

Best practices

  1. Build ThreadPoolExecutor explicitly with a bounded queue + a chosen rejection policy; avoid the Executors unbounded factories in production.
  2. Name your threads via ThreadFactory — future-you reading a thread dump will thank you.
  3. Size CPU pools to cores; size I/O pools with the wait/compute formula, then cap at the real downstream bottleneck.
  4. Always shutdown() + awaitTermination() + shutdownNow() on the way down.
  5. Treat InterruptedException as a first-class signal: propagate or restore-and-exit, never swallow.
  6. remove() every ThreadLocal in a finally when threads are pooled.
  7. Prefer CompletableFuture/structured concurrency over chains of blocking get().
  8. On Java 21+, use virtual threads for I/O-bound fan-out and stop hand-tuning pool sizes.

Interview Questions

Now let's practice everything you've learned in the shape of real interview questions. Answer each yourself first, then read the answer.

1. What's the difference between RUNNABLE and BLOCKED? Why does a socket read show as RUNNABLE?

BLOCKED means the thread is contending specifically for a synchronized monitor lock. WAITING/TIMED_WAITING mean it's parked awaiting a signal. RUNNABLE means the JVM considers it able to run. The JVM cannot see OS-level I/O blocking, so a thread stuck in a blocking socket read is still labeled RUNNABLE — the block happens below the JVM's visibility.

2. What happens if you call start() twice? What about calling run() directly?

A second start() throws IllegalThreadStateException — a thread is one-shot. Calling run() executes the body synchronously on the current thread; no new thread is created. This is a common accidental bug.

3. [HARD] You configure a ThreadPoolExecutor with core=4, max=50, and a newFixedThreadPool-style unbounded queue. Under heavy load, how many threads run?

Only 4. With an unbounded queue, tasks are always enqueued successfully, so the pool never reaches the "queue full" condition that triggers creating threads above core. maximumPoolSize is effectively dead. To use max, you need a bounded queue.

4. Walk through the exact submission algorithm of ThreadPoolExecutor.

(1) If threads < core, create a new thread. (2) Else try to enqueue. (3) If enqueue fails (queue full) and threads < max, create a non-core thread. (4) Else invoke the rejection handler. The subtle point: enqueue is attempted before growing past core, which is why an unbounded queue caps you at core.

5. Compare the four rejection policies and say which gives backpressure.

AbortPolicy throws (fail-fast, default). DiscardPolicy drops silently. DiscardOldestPolicy drops the head of the queue and retries. CallerRunsPolicy runs the task on the submitting thread — this provides backpressure because the submitter is busy and can't submit more, throttling intake.

6. Derive the thread count for: 16 cores, tasks that wait 190 ms on a downstream service and compute for 10 ms, target 100% utilization.

threads = cores × util × (1 + wait/compute) = 16 × 1.0 × (1 + 190/10) = 16 × 20 = 320. But also cap at the downstream's real concurrency limit — 320 threads hammering a service that allows 50 concurrent just queues; size to the bottleneck.

7. [HARD] Why is swallowing InterruptedException a serious bug, and what are the two correct handlings?

The interrupt flag is cleared when the exception is thrown. If you catch it and do nothing, you've destroyed the only cancellation signal — the task becomes uninterruptible and shutdownNow can't stop it. Correct: either rethrow/propagate it, or call Thread.currentThread().interrupt() to restore the flag and then exit the work.

8. shutdown() vs shutdownNow() — what happens to queued and running tasks in each?

shutdown(): stops accepting new tasks, lets queued and running tasks finish. shutdownNow(): stops accepting, returns un-started queued tasks as a List<Runnable> (they never run), and interrupts running tasks (which only stops them if they honor interruption).

9. [HARD] Why does using ThreadLocal in a thread pool cause both a correctness bug and a memory leak?

Correctness: pool threads are reused, so a value set for one task persists into the next task on the same thread unless you remove() it — leaking one request's context (e.g., user identity) into another. Memory: ThreadLocalMap keys are weak but values are strong; a long-lived pool thread holds the value even after the ThreadLocal is gone, until the slot is reused. remove() in a finally fixes both.

10. What does this print?
ExecutorService pool = Executors.newSingleThreadExecutor();
Future<?> f = pool.submit(() -> { throw new RuntimeException("boom"); });
System.out.println("submitted");
// no f.get() call
pool.shutdown();

It prints submitted and nothing about "boom". Exceptions from submitted tasks are captured in the Future and surface only on f.get() (as ExecutionException). Without get(), the exception is silently swallowed. (Had you used execute(), it would reach the UncaughtExceptionHandler and print a stack trace.)

11. scheduleAtFixedRate vs scheduleWithFixedDelay when a run overshoots its period?

scheduleAtFixedRate targets a fixed period; if a run exceeds the period, the next run starts immediately after (bursting to catch up, never overlapping the same task). scheduleWithFixedDelay always inserts the full delay after completion, so runs never pile up. For polling that must not overlap or stampede, use fixed delay.

12. [HARD] A scheduleWithFixedDelay task stops running after a while with no error in logs. Why?

It threw an uncaught exception. A scheduled task that throws is silently cancelled and never scheduled again; the exception is trapped in the (ignored) Future. Wrap the task body in try/catch to keep it alive.

13. Why must setDaemon(true) be called before start(), and when would a daemon thread lose data?

Daemon status is immutable once the thread is alive — setting it after start() throws IllegalThreadStateException. A daemon loses data when it's the only thing running a critical operation and all user threads exit: the JVM terminates and abandons the daemon mid-operation, without guaranteeing finally/cleanup runs. Never put must-complete work on daemons.

14. [HARD] On Java 21, you move a service to virtual threads but throughput doesn't improve and CPU shows carrier threads mostly idle. What's the likely cause?

Pinning. The blocking I/O sits inside a synchronized block (or a native call), so the virtual thread can't unmount from its carrier while blocked. With few carriers, pinned virtual threads starve the pool. Fix: replace synchronized around the blocking call with a ReentrantLock, which lets the virtual thread unmount. (Also verify you're using per-task virtual threads, not pooling them.)

15. When are virtual threads the wrong tool?

For CPU-bound work — they don't add parallelism beyond the carrier count, which is bounded by cores. Compute-heavy tasks should use a platform-thread pool sized to cores. Virtual threads only win when threads spend time waiting.

Senior notes & advanced edge cases

You now hold the mental model of threads, Executors, pool sizing, cooperative cancellation, and virtual threads. Now we go one layer deeper — the layer you only learn after a few on-call nights and a couple of real post-mortems: the things that look correct in code and detonate under production load. This section repeats none of the earlier material; it only fills the gaps.

Roadmap for this section
  1. The Java Memory Model: the happens-before you get for free from an Executor, and why you don't need volatile to read a task's result.
  2. availableProcessors() lies inside containers — the biggest sizing mistake in Kubernetes.
  3. ForkJoinPool.commonPool() and parallel streams share one pool, and blocking in it starves the whole JVM.
  4. CompletableFuture done right — which thread runs the callback, join vs get, collecting results.
  5. Thread-starvation deadlock (a pool nested in itself).
  6. Observability and hooks, context/MDC propagation, queue types.
  7. Spring @Async traps.
  8. The real virtual-threads playbook: a Semaphore instead of sizing, pinning in modern JDKs, ScopedValue, structured concurrency.
  9. Seven hard senior interview questions with full answers.

The memory model: happens-before you get for free

A subtle question: a task on thread B computes a value and writes it into an ordinary field (no volatile); thread A reads that field after future.get(). Is A guaranteed to see the fresh value, or might it see a stale/half-built one?

JMM guarantees for Executors

The java.util.concurrent package spec explicitly guarantees two happens-before edges:

  1. Everything you did on the submitting thread before submit/execute happens-before the task begins. The task sees all your pre-submit writes.
  2. Everything the task does happens-before Future.get() returns. After get(), thread A is guaranteed to see all of the task's writes, fully. So you do not need volatile or synchronized to hand a task's result across threads — the submit→run→get boundary establishes the visibility itself. The same guarantee holds for BlockingQueue (put/take), CountDownLatch, and a thread's termination relative to join().
Trap: fire-and-forget with no synchronization point

If you execute() a task and then read its result from another thread with no synchronization mechanism at all (no get, no queue, no latch, no volatile field), you have no happens-before edge and may see a stale value — or a half-constructed object — forever. The bug is non-deterministic and invisible in tests.

availableProcessors() lies inside containers

Every sizing formula from the previous chapter leans on "the number of cores." But where does that number come from? Runtime.getRuntime().availableProcessors(). The problem: in Kubernetes that number is often not what you think.

The biggest sizing mistake of recent years

Your node has 32 cores, but you gave the pod a cpu limit of, say, 500m (half a core). Since JDK 10 the JVM is container-aware and reads the cgroup; with a sub-core quota, availableProcessors() returns 1. Now three disasters:

  • If you sized a pool at a hard-coded 32, those 32 threads context-switch endlessly over a one-core quota and you get throttled.
  • ForkJoinPool.commonPool(), whose parallelism is availableProcessors()-1, effectively becomes 0 (one thread), so every parallel stream silently runs serially.
  • Conversely, without a limit but with an old/misconfigured JVM, it may see 32 and trample every neighbor on the node.
What to do

Log the real number at startup (log.info("cores={}", Runtime.getRuntime().availableProcessors())) — that one line erases hours of debugging. If a fractional limit forces a wrong count, set -XX:ActiveProcessorCount=N explicitly, and prefer cpu request == cpu limit for predictable behavior. (cgroups v2 is fully supported only from JDK 17 onward.)

commonPool() and parallel streams share one pool

stream().parallel() and CompletableFuture.supplyAsync(task) (the no-executor overload) both run on one global pool: ForkJoinPool.commonPool(). That pool is small (core-count) and built for short CPU-bound work.

commonPool starvation: one blocking call takes down the whole app

If you put a blocking call (HTTP, DB, Thread.sleep) inside a parallelStream().map(...), that worker holds a commonPool thread hostage. Because this pool is shared across the entire JVM, every other part of the app that relies on a parallel stream or supplyAsync now starves, and unrelated requests slow down. The classic incident: one reporting endpoint using a parallel stream drags the whole service to its knees.

Two correct fixes:

// Fix 1: run the parallel stream inside a dedicated ForkJoinPool
ForkJoinPool custom = new ForkJoinPool(16);
List<R> out = custom.submit(() ->
    items.parallelStream().map(this::blockingCall).toList()
).get();

// Fix 2 (for CompletableFuture): always pass your own executor
CompletableFuture.supplyAsync(this::blockingCall, myIoPool);
ManagedBlocker: the principled way to block inside a ForkJoinPool

If you must block inside an FJP, use ForkJoinPool.managedBlock(...) instead of a plain block. This interface tells the pool "I'm about to block," and the pool temporarily spins up a compensating worker to preserve parallelism. CompletableFuture itself uses this ManagedBlocker mechanism for join() when called on an FJP worker thread.

CompletableFuture done right: where the callback runs

The chapter introduced CompletableFuture as "non-blocking and composable." The senior point is one question: which thread runs the callback?

  • The non-Async form (e.g. thenApply): the callback runs on whichever thread completed the previous stage — or even on the calling thread if the stage was already complete. Unpredictable and dangerous for heavy work.
  • The Async form with no executor (thenApplyAsync): on commonPool() — with the starvation trap above.
  • The Async form with an executor: on your own pool. This is the correct production default.
CompletableFuture
    .supplyAsync(() -> fetchUser(id), ioPool)     // I/O on our pool
    .thenApplyAsync(this::enrich, cpuPool)         // CPU on a separate pool
    .exceptionally(ex -> User.anonymous())         // fallback
    .thenAccept(this::send);
`thenApply` vs `thenCompose`, `join` vs `get`, and error handling

If your function itself returns a CompletableFuture, thenApply leaves you with a nested CompletableFuture<CompletableFuture<T>>; use thenCompose (the flatMap of futures) to flatten. get() throws checked exceptions; join() is the unchecked version (CompletionException) and is cleaner inside chains. To gather many futures use allOf(a,b,c).thenApply(...) — but allOf returns Void, so join each one afterward. For errors: exceptionally catches only the error path, handle catches both, and whenComplete is a finally that observes but does not transform.

Thread-starvation deadlock: a pool nested in itself

This is one of the sneakiest deadlocks and never shows up in tests, because it only occurs under load.

The sequence: your pool is bounded (say 10 threads). A task running on that pool submits subtasks to the same pool and then get()s, waiting on their results. If 10 parent tasks simultaneously occupy all 10 threads and each waits on its own subtask, no thread is left to run any subtask — everyone waits on everyone; total lock.

Diagram — why a bounded, self-nested pool deadlocks:

sequenceDiagram
    participant P as Parent tasks (fill all N threads)
    participant Q as Work queue
    participant C as Child tasks
    P->>Q: submit child, then block on get()
    Q-->>C: waiting for a free thread
    Note over P,C: all N threads held by parents<br/>no thread free to run children
    Note over P,C: parents wait children, children wait threads = deadlock
Golden rule

A task running on a pool must never submit work to the same pool and then block on its result. Either use separate pools for separate layers, write the whole chain non-blocking with thenCompose, or use virtual threads (where threads are not scarce, which eliminates this whole class of deadlock).

Observability and hooks

A pool you can't see will ambush you one day. Feed getActiveCount(), getQueue().size(), and getCompletedTaskCount() to Micrometer — a perpetually full queue means the pool is too small or the downstream is slow. And override ThreadPoolExecutor's beforeExecute/afterExecute: the perfect place to set/clear MDC, time execution, and catch the exceptions that submit silently swallows (afterExecute receives both the Runnable and the Throwable). Two more knobs: allowCoreThreadTimeOut(true) lets idle core threads be reclaimed (low-traffic pools), and prestartAllCoreThreads() warms the lazily-created core threads up front when first-request latency matters.

Context/MDC propagation in async

SLF4J MDC ThreadLocals, SecurityContextHolder, and tracing trace-ids are not automatically carried onto pool threads — the worker doesn't have them. Fix: a TaskDecorator (in Spring) or a wrapper that copies the context from the submitting thread before running and clears it after. Virtual threads ease this pain but don't remove it.

Choosing the queue type is an engineering decision too: ArrayBlockingQueue is the safe production default, PriorityBlockingQueue for a prioritized queue, and DelayQueue for scheduling — all in preference to the unbounded, dangerous LinkedBlockingQueue.

Spring @Async traps

@Async is the prettiest "go async" in the world — until one of these bites you:

Three classic `@Async` traps
  1. Self-invocation: @Async works via a proxy. If a public method calls the @Async method on the same bean, the call never crosses the proxy and runs completely synchronously — with no error. You must invoke it from a different bean.
  2. Lost exceptions: if the @Async method returns void and throws, the exception goes nowhere (just like execute without a Future) and is lost from the logs. To catch it, define an AsyncUncaughtExceptionHandler, or make the return type a CompletableFuture so the error surfaces on get.
  3. The default executor: in plain Spring (no Boot), if you don't specify an executor the fallback is SimpleAsyncTaskExecutor, which has no pool and spins up a fresh thread for every call — thread explosion under load. Spring Boot mercifully auto-configures a ThreadPoolTaskExecutor, but don't trust even that default — configure it explicitly.
Spring Boot and virtual threads (the modern take)

Since Spring Boot 3.2, setting spring.threads.virtual.enabled=true (on Java 21+) switches both the @Async executor and Tomcat's request threads to virtual threads, so your traditional blocking code scales without a rewrite. But the same MDC and pinning caveats still apply.

The real virtual-threads playbook

The chapter said that with virtual threads you no longer size a pool. The senior's next question: so how do I protect a constrained downstream?

A Semaphore is the replacement for "pool size"

When you spawn one virtual thread per task, the pool is no longer a boundary — but your DB may allow only 20 connections. If 10,000 virtual threads hit it at once, the connection pool blows up or times out. The fix: bound concurrency not with pool size, but with a Semaphore around the scarce resource:

Semaphore db = new Semaphore(20);
db.acquire();
try { return jdbc.query(...); }
finally { db.release(); }

In the virtual-threads world, the Semaphore plays exactly the role maximumPoolSize used to play: a guard on downstream capacity.

ThreadLocal and expensive objects in a million-thread world

With platform threads a common trick was to cache an expensive object (a big buffer, say) in a ThreadLocal, because threads were few. With millions of virtual threads this is an anti-pattern: one copy per thread means explosive memory use. With virtual threads, either build the object per-task, pool it, or consider ScopedValue.

Modern update — pinning is nearly solved

The chapter said synchronized causes pinning. JEP 491 in JDK 24 fixed exactly that: a synchronized block no longer normally holds the carrier hostage. Only native/JNI and a few rare corners remain. To monitor pinning, use the JFR event jdk.VirtualThreadPinned instead of the old -Djdk.tracePinnedThreads flag (deprecated in JDK 24). You tune carrier count with -Djdk.virtualThreadScheduler.parallelism.

Build one manually with Thread.ofVirtual().name("job-", 0).start(runnable) or Thread.startVirtualThread(r). Two modern companions round out the fan-out pattern: structured concurrency (still a preview in JDK 25, JEP 505, now via StructuredTaskScope.open() with a Joiner) launches several subtasks in one scope so they all fail/cancel together and never leak a thread; and ScopedValue (finalized in JDK 25, JEP 506) is the immutable, cheaper ThreadLocal replacement for passing context to subtasks — no remove() needed, since its scope closes automatically.

Hard senior interview questions

1. [HARD] A task on another thread writes a value into an ordinary (non-volatile) field, and you read it after `future.get()`. Do you need volatile?

No. The JMM guarantees that everything the task does happens-before Future.get() returns, and everything you wrote before submit happens-before the task starts. So the submit→run→get boundary establishes full visibility; no volatile/synchronized is needed to hand the result across. But if you read the result from a fire-and-forget task with no synchronization point (no get, no queue, no latch), you have no guarantee and a stale value is possible.

2. [HARD] Your k8s pod runs with `cpu limit: 1` on a 32-core node. `newFixedThreadPool(Runtime.getRuntime().availableProcessors()*2)` and parallel streams behave oddly. Why?

Because since JDK 10 the JVM is container-aware and reads the cgroup: with a one-core limit, availableProcessors() returns 1, not 32. So your pool is 2 threads, and worse, commonPool() (parallelism cores-1) becomes single-threaded, so parallel streams run serially. Fix: log the real number, set -XX:ActiveProcessorCount explicitly if needed, and prefer cpu request == limit.

3. [HARD] An endpoint using `parallelStream().map(remoteCall)` slows down completely unrelated requests. Why, and two fixes?

Because a parallel stream runs on the global ForkJoinPool.commonPool(), and the blocking remote call holds its threads hostage; since that pool is shared JVM-wide, every other user of it starves. Fix 1: submit the stream inside a dedicated ForkJoinPool. Fix 2 (better): don't use a parallel stream for blocking work — use CompletableFuture.supplyAsync(task, dedicatedIoPool) or virtual threads. If you must block inside an FJP, use a ManagedBlocker so the pool spins up a compensating worker.

4. [HARD] You have a bounded 10-thread pool; each task submits subtasks to the same pool and blocks on `get`. Under load the app hangs. Diagnose it.

Thread-starvation deadlock. If 10 parent tasks simultaneously hold all 10 threads and each waits on its own subtask, no thread is left to run any subtask — parents wait on children, children wait on a free thread. Fix: separate pools per layer, a non-blocking chain with thenCompose, or virtual threads, where threads aren't scarce.

5. `thenApply` vs `thenApplyAsync` — where does each callback run, and where does it bite? Difference between `get` and `join`?

thenApply: the callback runs on whichever thread completed the previous stage (or even the calling thread if already complete) — unpredictable for heavy work. thenApplyAsync with no executor: on commonPool (the starvation trap). thenApplyAsync with an executor: on your pool — the correct production form. get() throws checked exceptions (ExecutionException); join() is unchecked (CompletionException) and cleaner inside chains.

6. [HARD] You call an `@Async` method from within the same class but it runs synchronously. Why? And why do `@Async void` exceptions vanish?

Because @Async works via a proxy; a self-invocation never crosses the proxy, so no async wrapping happens and the code runs synchronously — you must call it from another bean or self-inject the proxy. On exceptions: an @Async method returning void is like execute without a Future; the exception is caught nowhere and lost from the logs. Fix: define an AsyncUncaughtExceptionHandler or return a CompletableFuture so the error surfaces on join/get.

7. [HARD] With virtual threads you no longer size a pool; so how do you protect a downstream that allows only 20 concurrent connections? And how do you detect pinning now?

You bound concurrency with a Semaphore(20) around the downstream call; in the virtual-threads world the Semaphore plays the role maximumPoolSize used to. For pinning: in JDK 24, JEP 491 means synchronized usually no longer pins; only native/JNI remains. To monitor, use the JFR event jdk.VirtualThreadPinned (the old -Djdk.tracePinnedThreads flag was deprecated in JDK 24).

Senior recap
  • submit → run → Future.get establishes happens-before; no volatile needed to hand a result across, but fire-and-forget with no sync point is unsafe.
  • Inside a container, availableProcessors() comes from the cgroup and may be far below the node's cores — breaking every sizing formula.
  • commonPool() and parallel streams are global; blocking in them starves the whole JVM. Give a dedicated executor or use a ManagedBlocker. And a bounded self-nested pool is a thread-starvation deadlock.
  • In CompletableFuture use the Async form with your own executor; make the pool observable and propagate MDC with a decorator.
  • @Async traps: self-invocation goes synchronous, void swallows exceptions, the default may have no pool.
  • With virtual threads: a Semaphore replaces sizing, no expensive ThreadLocals, pinning is nearly solved in JDK 24; ScopedValue and structured concurrency are the modern fan-out pattern.
In a nutshell
  • A platform thread is expensive (~1 MB stack + a system call); that's why Executors decouple task submission from thread creation.
  • Memorize the six Thread.State values. BLOCKED is only monitor contention; OS-level I/O shows as RUNNABLE.
  • run() creates no thread, start() does. Runnable returns nothing, Callable does and you fetch it with Future.get() — which is blocking and hides exceptions until the moment of get().
  • In production, build ThreadPoolExecutor by hand with a bounded queue and a rejection policy. With an unbounded queue, maximumPoolSize is inert.
  • Size CPU-bound to core count; size I/O-bound with cores × (1 + wait/compute), and cap at the real bottleneck.
  • For scheduling, know fixed-rate vs fixed-delay and wrap the task body in try/catch (else silent death).
  • Always shutdown()awaitTermination()shutdownNow(). Cancellation is cooperative; never swallow InterruptedException.
  • In pools, remove() every ThreadLocal in a finally (correctness + security + memory).
  • On Java 21+, for I/O-bound work create one virtual thread per task and forget sizing — just watch out for pinning and don't pool them.

Sources: ThreadPoolExecutor (JDK 17), Managing Throughput with Virtual Threads — inside.java, RejectedExecutionHandler — Baeldung