Concurrency · همزمانی متوسطIntermediate ~68 دقیقه مطالعه~59 min read
نخها، Runnable/Callable و ExecutorهاThreads, Runnable/Callable & Executors
از صفر تا سنیور یاد میگیری که نخ چیست، چطور با Runnable/Callable/Future کار بسپاری، چرا Executorها را دوست داری، ThreadPoolExecutor را چطور بسازی و اندازه بگیری، چطور باوقار خاموش کنی، لغو تعاونی و ThreadLocal را رام کنی، و کجا نخهای مجازی همهچیز را عوض میکنند.A zero-to-senior walkthrough of what a thread really is, how to hand off work with Runnable/Callable/Future, why Executors exist, how to build and size a ThreadPoolExecutor, shut it down gracefully, tame cooperative cancellation and ThreadLocal, and where virtual threads rewrite the rules.
خب، بنشین کنار من. میخواهیم دربارهٔ یکی از آن موضوعهایی حرف بزنیم که همه در مصاحبه ازش میترسند و در تولید (production) ازش خوندل میخورند: همزمانی (concurrency). اما قول میدهم اگر سه چیز را درست بفهمی، این موضوع از حالت «جادوی ترسناک» درمیآید و میشود چیزی که کف دستت است. تمام این فصل، همان سه چیز است که با آجرهای کوچک، یکییکی، از پایه میسازیمشان.
قبل از هر چیز یک تصویر ذهنی بگیر: یک نخ (thread) مثل یک کارگر است که میتواند یک کار را دنبال کند. اگر یک کارگر داشته باشی، کارها پشت سر هم انجام میشوند. اگر چند کارگر داشته باشی، کارها میتوانند همزمان پیش بروند. تمام دعوای این فصل سر این است که این کارگرها را چطور بسازی، چطور بهشان کار بدهی، و چطور مؤدبانه بگویی «دیگر بس است، برو خانه».
در این فصل این مسیر را میرویم:
- نخ واقعاً چیست و چرا ساختنش گران است (و همین گرانی، دلیل وجود Executorهاست).
- چرخهٔ حیات نخ — شش حالتی که یک نخ میتواند در آن باشد.
- واحد کار:
Runnable،Callableو دستگیرهٔ نتیجه یعنیFuture. - چارچوب Executor و قلبش
ThreadPoolExecutorبا هفت پارامترش. - ریاضیِ اندازهگذاری استخر — همان جایی که سنیورها امتیاز میگیرند.
- زمانبندی (
ScheduledExecutorService) و خاموشسازیِ باوقار. - لغو تعاونی (
InterruptedException) وThreadLocalکه در استخرها نشت میکند. - نخهای daemon و در پایان نخهای مجازی (Java 21+) که کل ریاضیِ اندازهگذاری را دود میکنند. در انتها پانزده پرسش مصاحبه با پاسخ کامل داریم. آمادهای؟ برویم.
بخش ۰ — واژههایی که باید از پیش بشناسی
قبل از اینکه بپریم وسط کد، بیا چند واژه را که مدام به کار میبریم، با یک تشبیه کوچک جا بیندازیم. اگر اینها ته ذهنت باشد، بقیهٔ فصل مثل آب خوردن میشود.
یک رستوران را تصور کن. آشپز همان نخ است: یک نفر که میتواند یک سفارش را از اول تا آخر بپزد. مدیر سالن که تصمیم میگیرد کدام آشپز الان کار کند همان زمانبند سیستمعامل (OS scheduler) است — چون آشپزخانه فقط چند اجاق (CPU/هسته) دارد و همه نمیتوانند همزمان بپزند، مدیر نوبت میدهد. میز کار شخصی هر آشپز با تخته و چاقو و ادویهاش، همان پشته (stack) است: حافظهای که فقط مال خودِ آن نخ است. رفتن تا انبار و برگشتن برای گرفتن یک آشپز جدید، همان فراخوان سیستمی (system call) است: کند و پرهزینه. برای همین رستورانهای حرفهای آشپز را وسط شیفت استخدام و اخراج نمیکنند؛ یک تیم ثابت دارند و سفارشها را بینشان پخش میکنند. این دقیقاً همان چیزی است که Executor برایت انجام میدهد.
- نخ (thread): واحد اجرای مستقل. یک نخ یک دنباله از دستورها را جلو میبرد.
- زمانبند سیستمعامل (OS scheduler): بخشی از سیستمعامل که تصمیم میگیرد کدام نخ روی کدام هسته و چه مدت اجرا شود. تو مستقیم کنترلش نمیکنی، فقط درخواست میکنی.
- پشته (stack): حافظهٔ خصوصیِ هر نخ برای متغیرهای محلی و زنجیرهٔ فراخوان متدها. هر نخ پشتهٔ خودش را دارد (معمولاً حدود ۱ مگابایت رزرو).
- فراخوان سیستمی (system call): وقتی برنامهٔ تو از هستهٔ سیستمعامل چیزی میخواهد (مثلاً «یک نخ تازه بساز»). اینها نسبتاً کندند.
- مسدود شدن (blocking): وقتی یک نخ میایستد و منتظر میماند تا اتفاقی بیفتد — مثلاً منتظر پاسخ دیتابیس. در این مدت آن نخ کاری نمیکند جز انتظار.
- قفل مانیتور (monitor lock): قفلی که با کلیدواژهٔ
synchronizedروی یک شیء گرفته میشود تا فقط یک نخ همزمان واردِ ناحیهٔ حساس شود. مثل کلید دستشویی که فقط یکی دستش است. - CPU-محور در برابر I/O-محور: کارِ CPU-محور یعنی نخ واقعاً محاسبه میکند (فشردهسازی، هش). کارِ I/O-محور یعنی نخ بیشتر منتظر چیزی بیرون است (دیتابیس، شبکه، دیسک). این تمایز، ستون فقرات اندازهگذاری استخر است.
خب، حالا که ابزارها دستمان است، برویم سراغ خودِ نخ.
مدل ذهنی: نخ چیست و چرا گران است
یک java.lang.Thread در جاوا در واقع پوششی نازک روی یک نخ سیستمعامل است — به این نخِ سیستمعاملی میگوییم نخ سکویی (platform thread)، تا بعداً از نخ مجازی جدایش کنیم.
حالا چرا میگویم گران است؟ وقتی یک نخ میسازی، سه اتفاق پرهزینه میافتد:
- یک پشتهٔ بومی بزرگ تخصیص داده میشود (معمولاً حدود ۱ مگابایت رزرو).
- نخ نزد زمانبند سیستمعامل ثبت میشود.
- همهٔ اینها یک فراخوان سیستمی میطلبد.
همین گرانیِ ساختِ نخ، تمام دلیل وجود چارچوب Executor است. بهجای اینکه برای هر کار یک آشپز تازه استخدام کنی (کُند و گران)، یک تیم کوچکِ ثابت از آشپزها داری و کارها را بینشان پخش میکنی. به این میگویند جدا کردن ثبت وظیفه از ساخت نخ (decouple task submission from thread creation). کار را میسپاری؛ اینکه کدام نخِ آماده اجرایش کند، دیگر دغدغهٔ تو نیست.
سه پرسش کل این فصل را قاب میگیرند. اگر جواب این سه را بلد باشی، همزمانی دیگر رازآلود نیست:
- واحد کار چیست؟
Runnable(بدون نتیجه، بدون استثنای checked)،Callable<V>(مقدار برمیگرداند و میتواند استثنای checked پرتاب کند)، یا زیرکلاس خامِThread(که تقریباً هیچوقت پاسخ درست نیست). - چه کسی و روی کدام نخ اجرایش میکند؟ یک
Threadخام، یا یکExecutorServiceکه استخری (pool) پشتش است. - چطور متوقف میشود؟ بهصورت تعاونی، از طریق وقفه (interruption) — چون جاوا هیچ
Thread.stop()امنِ پیشدستانهای (preemptive) ندارد.
بیا همین سه را یکییکی باز کنیم.
چرخهٔ حیات و حالتهای نخ
هر نخ در طول عمرش بین چند «حالت» جابهجا میشود. جاوا این حالتها را در enumی به نام Thread.State نگه میدارد. این enum حقیقتِ پایه است — حفظش کن، چون مصاحبهگر تقریباً همیشه تفاوتِ WAITING و BLOCKED را میکاود.
یک کارمند را تصور کن. NEW: تازه استخدام شده ولی هنوز سرِ کار نیامده. RUNNABLE: پشت میزش نشسته و کار میکند یا آمادهٔ کار است. BLOCKED: پشت درِ اتاق کنفرانس منتظر است تا نفر قبلی بیرون بیاید و کلید (قفل) را بدهد. WAITING: منتظر است تا همکارش بهش خبر بدهد، بدون هیچ مهلتی. TIMED_WAITING: همان انتظار، ولی با ساعت زنگدار — «تا ساعت ۳ صبر میکنم». TERMINATED: کارش تمام شده و رفته خانه.
| حالت | معنا | چطور به آن میرسی |
|---|---|---|
NEW |
ساختهشده، هنوز start نشده | new Thread(...) پیش از start() |
RUNNABLE |
آمادهٔ اجرا یا در حال اجرا (JVM «آماده» و «روی CPU» را تفکیک نمیکند) | بعد از start() |
BLOCKED |
منتظر گرفتنِ قفل مانیتور (synchronized) |
رقابت بر سرِ یک مانیتور |
WAITING |
انتظار نامحدود برای سیگنالِ نخ دیگر | wait(), join(), LockSupport.park() |
TIMED_WAITING |
همان، ولی با مهلت | sleep(ms), wait(ms), join(ms), park(ns) |
TERMINATED |
run() بازگشت یا پرتاب کرد |
وظیفه پایان یافت |
حالا دو تلّه که دقیقاً همان جایی است که مصاحبهگر انگشت میگذارد:
این یکی خیلیها را غافلگیر میکند. نخی که روی یک عملیات I/O مسدود است (مثلاً وسط خواندن از یک سوکت شبکه منتظر داده مانده) در thread dump بهصورت RUNNABLE نشان داده میشود، نه BLOCKED. چرا؟ چون این انسداد در سطح سیستمعامل رخ میدهد و JVM اصلاً آن را نمیبیند — از دید JVM آن نخ «قادر به اجرا» است. برچسب BLOCKED را جاوا فقط و فقط برای رقابت بر سرِ قفلِ مانیتورِ synchronized استفاده میکند. پس دفعهٔ بعد که thread dump دیدی و همهجا RUNNABLE بود، فریب نخور؛ ممکن است نخها در واقع منتظر شبکه باشند.
start() یکبارمصرف است. اگر start() را دوبار صدا بزنی، استثنای IllegalThreadStateException میگیری. و مهمتر: اگر مستقیم run() را صدا بزنی، اصلاً نخِ جدیدی راه نمیافتد — بدنه روی نخِ فعلی و بهصورت همریسمانی (synchronous، یعنی همان لحظه و همانجا) اجرا میشود. یعنی هیچ همزمانیای اتفاق نمیافتد. این یکی از رایجترین باگهای تصادفیِ تازهکارهاست.
Runnable task = () -> System.out.println("on " + Thread.currentThread());
Thread t = new Thread(task);
t.run(); // باگ: روی نخِ *فعلی* اجرا میشود، بدون همزمانی
t.start(); // درست: نخ جدید میسازد
فرق start() و run() را اینطور در ذهنت نگه دار: run() فقط یک متد عادی است که تو صدایش میزنی؛ start() است که به سیستمعامل میگوید «یک نخ تازه بساز و آن نخ، run() را صدا بزند».
Thread در برابر Runnable در برابر Callable/Future
حالا برویم سراغ پرسش اول: واحد کار چیست؟
یک قاعدهٔ طلایی: ترکیب (composition) را بر وراثت (inheritance) ترجیح بده. یعنی بهجای اینکه کلاست را از Thread ارثبری کنی (extends Thread)، کارت را بهشکل یک Runnable یا Callable بنویس و آن را بده به نخ. چرا؟ چون وقتی Thread را زیرکلاس میکنی، کارت را برای همیشه به یک نخ گره میزنی و دیگر نمیتوانی همان کار را در یک استخر بازاستفاده کنی. اما یک Runnable یک بستهٔ کار مستقل است که هر نخی میتواند بازش کند و انجامش دهد.
Runnable/Callable مثل یک برگهٔ سفارش است: روی کاغذ نوشتهای «یک پیتزا بپز». این برگه را میشود به هر آشپزی داد. اما extends Thread مثل این است که برای هر سفارش یک آشپز مخصوص همان سفارش بسازی که کارش توی وجودش دوخته شده و به درد سفارش بعدی نمیخورد. برگهٔ سفارش قابل بازاستفاده است؛ آشپزِ دوختهشده نه.
تفاوت Runnable و Callable:
Runnableیعنی «بینداز و فراموش کن»: نه مقداری برمیگرداند، نه در امضایش اجازهٔ استثنای checked دارد.Callable<V>یعنی «کاری کن و نتیجهاش را به من برگردان»: یک مقدار از نوعVبرمیگرداند و میتواند استثنای checked پرتاب کند.
// Runnable: بینداز و فراموش کن، بدون بازگشت، بدون checked در امضا
Runnable r = () -> log.info("done");
// Callable: مقدار برمیگرداند، میتواند checked پرتاب کند
Callable<Integer> c = () -> {
return httpClient.fetchCount(); // ممکن است IOException بدهد
};
خب، اگر Callable نتیجه برمیگرداند، اما همین حالا اجرا نمیشود و بعداً روی نخِ دیگری تمام میشود — چطور نتیجهاش را بگیرم؟ اینجاست که Future وارد میشود.
Future<V> مثل فیشی است که خشکشویی به تو میدهد. کارت (لباس) را سپردهای، الان نتیجه آماده نیست، ولی یک دستگیره داری که با آن بعداً میتوانی نتیجه را تحویل بگیری. وقتی f.get() را صدا میزنی، مثل این است که فیش را ببری و بگویی «حالا لباسم را بده» — و اگر هنوز آماده نباشد، همانجا میایستی و منتظر میمانی (این یعنی get() مسدودکننده است).
ExecutorService pool = Executors.newFixedThreadPool(4);
Future<Integer> f = pool.submit(c);
try {
Integer n = f.get(2, TimeUnit.SECONDS); // تا ۲ ثانیه مسدود میشود
} catch (ExecutionException e) {
// وظیفه پرتاب کرد؛ e.getCause() استثنای اصلی است
Throwable real = e.getCause();
} catch (TimeoutException e) {
f.cancel(true); // true => نخِ در حال اجرا را interrupt کن
}
سه نکته دربارهٔ Future که مردم را میلغزاند — اینها را خوب بخوان چون هر سه میتوانند در تولید باگهای خاموش بسازند:
این خطرناکترین رفتار Future است. اگر وظیفهای که با submit() سپردهای وسط کار استثنا پرتاب کند، هیچ چیز بهظاهر اتفاق نمیافتد. نه لاگی، نه خطایی، هیچ. استثنا در یک ExecutionException پیچیده میشود و فقط و فقط وقتی get() را صدا بزنی به تو تحویل داده میشود (و استثنای اصلی در e.getCause() است). در مقابل، اگر با execute(Runnable) — که اصلاً Future نمیدهد — کار بسپاری، استثنای گرفتهنشده به UncaughtExceptionHandler آن نخ میرود و معمولاً چاپ میشود. این عدمتقارن، باگها را بیصدا دفن میکند. قاعده: یا همیشه روی Futureهایت get() بزن، یا آگاهانه از execute استفاده کن.
f.cancel(true)فقط پرچم وقفه را ست میکند؛ خودِ وظیفه باید همکاری کند (پایینتر در بخش وقفه توضیحش میدهم).cancel(false)جلوی اجرای وظیفهای را که هنوز شروع نشده میگیرد، اما وظیفهٔ در حال اجرا را interrupt نمیکند.get()مسدودکننده است. یعنی نخِ فراخوان همانجا میایستد تا نتیجه آماده شود. اگر میخواهی کارها را زنجیر کنی بدون اینکه بایستی،CompletableFutureرا ترجیح بده که نامسدودکننده و قابلترکیب است (با متدهایی مثلthenApply,thenCompose,allOf).
چارچوب Executor
حالا پرسش دوم: چه کسی کار را اجرا میکند؟ جوابِ حرفهای: یک Executor.
چارچوب Executor یک سلسلهمراتب تمیز دارد که از ساده به کامل میرود:
Executor— پایهترین، فقط یک متد دارد:execute(Runnable).ExecutorService— این را گسترش میدهد وsubmit،invokeAllو مدیریتِ چرخهٔ حیات (خاموشسازی) را اضافه میکند.ScheduledExecutorService— این هم آن را گسترش میدهد و تأخیر و اجرای دورهای را اضافه میکند.
اما اسبِکارِ واقعی که همهٔ اینها پشتش هست، کلاسِ ThreadPoolExecutor است. بیا اول ببینیم چرا نباید کورکورانه به میانبُرهای آماده اعتماد کنی.
کارخانهٔ Executors — و چرا باید بهش بدگمان باشی
کلاس Executors چند متدِ کارخانهایِ (factory) راحت دارد که در یک خط برایت استخر میسازند. مشکل اینجاست که چندتایشان در محیط تولید خطرناکاند، چون یا از صفهای بیکران (unbounded queue) استفاده میکنند یا تعداد نخِ بیکران:
Executors.newFixedThreadPool(n); // LinkedBlockingQueue بدون کران ظرفیت → خطر OOM
Executors.newSingleThreadExecutor();// همان صف بیکران
Executors.newCachedThreadPool(); // maxPoolSize = Integer.MAX_VALUE → نخ بیکران
بگذار خطر را ملموس کنم:
«بیکران» یعنی «بدون سقف». زیر بارِ مداوم و سنگین، newFixedThreadPool صفِ کارهای منتظرش را بیحد رشد میدهد — کار روی کار انباشته میشود تا هیپ (heap) پر شود و OutOfMemoryError بگیری. از آن طرف، newCachedThreadPool سقفی روی تعداد نخ ندارد (Integer.MAX_VALUE)، پس زیر فشار آنقدر نخ میسازد تا سیستمعامل دیگر نخ نسازد و برنامه از پا دربیاید. رویهٔ سنیور یک جمله است: ThreadPoolExecutor را مستقیم و دستی بساز، با یک صفِ کراندار و یک سیاستِ ردّ (rejection policy) صریح. اینطور وقتی سیستم زیر فشار است، بهجای ترکیدن، مؤدبانه «نه» میگوید.
ThreadPoolExecutor: هفت پارامتر
بیا سازندهٔ کامل را ببینیم. هفت پارامتر دارد و هر کدام یک تصمیم مهندسی است:
ThreadPoolExecutor pool = new ThreadPoolExecutor(
4, // corePoolSize
16, // maximumPoolSize
60L, TimeUnit.SECONDS, // keepAliveTime برای نخهای بالای core
new ArrayBlockingQueue<>(1000), // صف کاری کراندار
new ThreadFactoryBuilder() // نخها را نامگذاری کن! (Guava یا کارخانهٔ دستساز)
.setNameFormat("api-worker-%d").build(),
new ThreadPoolExecutor.CallerRunsPolicy() // handlerِ ردّ
);
- corePoolSize (۴): تعداد نخهای همیشگی که حتی وقتی بیکارند نگه داشته میشوند.
- maximumPoolSize (۱۶): بیشترین تعداد نخی که استخر مجاز است تا آن بسازد.
- keepAliveTime (۶۰ ثانیه): نخهای اضافه بر core اگر اینقدر بیکار بمانند، از بین میروند.
- صف کاری (ArrayBlockingQueue با ظرفیت ۱۰۰۰): جایی که کارها منتظر میمانند تا نخی آزاد شود. اینکه این صف کراندار باشد، حیاتی است — همین الان میفهمی چرا.
- ThreadFactory: کارخانهای که نخها را میسازد؛ اینجا از آن برای نامگذاری نخها استفاده کردهایم.
- RejectedExecutionHandler: وقتی همهچیز پر است، با کارِ تازه چه کنیم.
تصور کن یک کافه داری. corePoolSize یعنی تعداد باریستاهای ثابت (مثلاً ۴). صف، یعنی صندلیهای انتظار (۱۰۰۰ تا). maximumPoolSize یعنی بیشترین باریستایی که میتوانی از انبار صدا بزنی (۱۶). حالا وقتی مشتری میآید: اول اگر باریستای ثابتِ آزاد کم داری، یکی سرِ کار میآید. بعد اگر همه مشغولاند، مشتری روی صندلی انتظار مینشیند. فقط وقتی همهٔ صندلیها هم پر شد، تازه باریستای اضافه از انبار صدا میزنی. و اگر باریستاها هم به سقف رسیدند و صندلیها هم پر بود، دمِ در به مشتری «نه» میگویی.
الگوریتم ثبتِ ThreadPoolExecutor دقیقاً همین ترتیب است و ارزش حفظکردن دارد، چون یک شگفتیِ بدنام را توضیح میدهد:
- اگر نخهای در حال اجرا <
corePoolSize، نخ جدید بساز برای وظیفه (حتی اگر نخهای بیکارِ دیگری وجود داشته باشند). - وگرنه، سعی کن وظیفه را در صف بگذاری.
- اگر صف پر است و نخهای در حال اجرا <
maximumPoolSize، نخ جدید (غیر-core) بساز. - اگر صف پر است و نخها به max رسیدهاند، از طریق handler ردّ کن.
حالا تلّه را ببین. اگر صفت بیکران باشد (مثل همان newFixedThreadPool)، گام ۲ همیشه موفق میشود — صف که هیچوقت پر نمیشود. یعنی شرطِ «صف پر است» در گام ۳ هرگز برقرار نمیشود، پس گام ۳ اصلاً فعال نمیشود و maximumPoolSize کاملاً بیاثر میماند. استخر هرگز از corePoolSize فراتر نمیرود، هر چقدر هم بار زیاد شود. این تقریباً همه را غافلگیر میکند. پس اگر میخواهی استخرت زیر بار بزرگ شود، حتماً باید صفِ کراندار بدهی (یا یک SynchronousQueue که ظرفیتش صفر است — پس هر وظیفه یا فوراً یک نخِ نو میطلبد یا ردّ میشود؛ newCachedThreadPool دقیقاً همینطور کار میکند).
سیاستهای ردّ (RejectedExecutionHandler)
وقتی استخر کاملاً اشباع است (صف پر و نخها در حداکثر)، با کارِ تازه چه کنیم؟ چهار سیاست آماده داری:
| سیاست | رفتار | چه زمانی |
|---|---|---|
AbortPolicy (پیشفرض) |
RejectedExecutionException پرتاب میکند |
میخواهی سریع شکست بخوری و فراخوان، سرریز را مدیریت میکند |
CallerRunsPolicy |
وظیفه را روی نخِ ثبتکننده اجرا میکند | فشار برگشتی (backpressure): ثبتکننده کند میشود چون خودش مشغول اجرای وظیفه است، پس ورودی بهطور طبیعی throttle میشود |
DiscardPolicy |
وظیفه را بیصدا میاندازد دور | فقط وقتی از دست دادن کار قابل قبول است (مثلاً متریکِ best-effort) |
DiscardOldestPolicy |
قدیمیترین وظیفهٔ صف را میاندازد و ثبت را دوباره امتحان میکند | بهندرت درست؛ کاری را میاندازد که بیشترین انتظار را کشیده |
«فشار برگشتی» یعنی وقتی سیستم زیر فشار است، بهجای اینکه کورکورانه بیشتر و بیشتر کار قبول کند تا بترکد، به سرچشمه فشار میآورد که «آهستهتر بفرست». CallerRunsPolicy این را رایگان به تو میدهد: وقتی استخر پر است، خودِ نخی که کار را سپرده مجبور میشود آن کار را اجرا کند. تا وقتی مشغول اجرای آن است، نمیتواند کار جدید بسپارد — پس نرخ ورود خودبهخود کم میشود. این پیشفرضِ عملگرایانه برای خطوط لولهٔ درخواست است. اما یک هشدار: اگر آن نخِ ثبتکننده همان نخی باشد که اتصالهای HTTP را میپذیرد، اجرای وظیفه روی آن یعنی تا مدتی پذیرشِ اتصال جدید متوقف میشود. که ممکن است خواستهٔ تو باشد یا نباشد.
ریاضیِ اندازهگذاری استخر
خب رسیدیم به جایی که نامزدهای سنیور واقعاً امتیاز میگیرند: چند نخ بگذارم؟ جواب به یک چیز بستگی دارد — کارت CPU-محور است یا I/O-محور.
حالت CPU-محور: کارِ CPU-محور (فشردهسازی، هش، محاسبهٔ محض) یک هسته را تماممدت مشغول نگه میدارد. اینجا اگر نخِ بیشتر از تعداد هستهها بسازی، هیچ سودی نمیبری؛ فقط سربارِ تعویضبافت (context switch — یعنی هزینهای که سیستمعامل میدهد تا نخی را کنار بگذارد و نخِ دیگری را روی هسته بنشاند) اضافه میکنی:
threads ≈ تعداد هستهها (یا cores + 1 برای پوششِ گاهبهگاهِ page fault)
اگر ۴ اجاق داری و ۴ آشپز، هر اجاق یک آشپز، همهچیز عالی است. حالا ۸ آشپز بیاور در همان ۴ اجاق. آیا سریعتر میپزی؟ نه! چون فقط ۴ اجاق هست. آشپزهای اضافه یا بیکار میایستند یا مدام سرِ اجاق با هم جا عوض میکنند و همین جابهجایی خودش وقت میبرد. کارِ CPU-محور یعنی «اجاق همیشه روشن است»؛ پس بیشترین سرعت وقتی است که تعداد آشپز = تعداد اجاق.
حالت I/O-محور: کارِ I/O-محور بیشترِ وقتش را منتظر میماند (منتظر دیتابیس، HTTP، دیسک). اینجا داستان کاملاً فرق دارد: میخواهی آنقدر نخ داشته باشی که وقتی چند نخ منتظر ماندهاند، بقیه بتوانند از CPU استفاده کنند. فرمول معروفِ برایان گوتز (از کتاب Java Concurrency in Practice)، که در واقع بازگوییِ قانون لیتل (Little's Law) است:
threads = cores × targetUtilization × (1 + waitTime / computeTime)
بیا با یک مثال حلشده جا بیندازیم. فرض کن ۸ هسته داری، هدفت ۱۰۰٪ بهرهوری است، و هر وظیفه ۹۰ میلیثانیه منتظر دیتابیس میماند و ۱۰ میلیثانیه محاسبه میکند (پس نسبت انتظار/محاسبه = ۹):
threads = 8 × 1.0 × (1 + 90/10) = 8 × 10 = 80 نخ
دیدی چه شد؟ یک سرویسِ I/O-سنگین بهدرستی حدود ۸۰ نخ روی ۸ هسته میخواهد — ده برابرِ تعداد هستهها! منطقش ساده است: چون هر نخ ۹۰٪ وقتش را فقط منتظر است، برای اینکه هستهها بیکار نمانند به نخهای زیادی نیاز داری تا همیشه یکیشان کار مفید کند. کم بگیری، CPU بیکار میماند و درخواستها صف میکشند؛ زیاد بگیری، حافظه میسوزانی (هر نخ ~۱ مگابایت پشته) و سیستم را به thrash (جابهجاییِ بیفایدهٔ مداوم) میاندازی. این نسبت انتظار/محاسبه، تکعددِ حیاتیِ اندازهگذاری استخر است — و دقیقاً همان فشاری است که نخهای مجازی (که آخر فصل میبینی) برطرفش میکنند.
int cores = Runtime.getRuntime().availableProcessors();
double waitToCompute = 9.0; // از پروفایلینگ اندازهگیری شده، نه حدس
int size = (int) (cores * (1 + waitToCompute));
یک نکتهٔ عملیِ مهم: عددِ waitToCompute را باید از پروفایلینگِ واقعی دربیاوری، نه از حدس.
فرمول یک عدد به تو میدهد، اما دنیای واقعی یک سقفِ سختِ دیگر هم دارد. اگر استخرِ اتصال دیتابیسات فقط ۲۰ اتصال دارد، بیش از ~۲۰ نخِ همزمانِ دیتابیس فایدهای ندارد — نخِ ۲۱ام فقط پشت استخرِ اتصال صف میکشد. پس همیشه استخرت را به اندازهٔ گلوگاه (bottleneck) واقعیِ پاییندستی کران بگذار. اندازهگیری کن که واقعاً کجا تنگ میشود و همانجا بایست.
ScheduledExecutorService
گاهی نمیخواهی کار همین حالا اجرا شود؛ میخواهی «۵ ثانیهٔ دیگر» یا «هر ۱۰ ثانیه یکبار» اجرا شود. برای اینها ScheduledExecutorService را داری.
شاید کلاس قدیمیِ Timer را بشناسی. در عمل منسوخ حساب میشود، به دو دلیل: (۱) فقط از یک نخ استفاده میکند، پس اگر یک وظیفه کند باشد بقیه عقب میافتند؛ و (۲) اگر یک وظیفه استثنا پرتاب کند، کلِ Timer بیصدا میمیرد. ScheduledExecutorService جایگزینِ درست و مقاومِ آن است.
ScheduledExecutorService sched = Executors.newScheduledThreadPool(2);
sched.schedule(() -> log.info("once, after 5s"), 5, TimeUnit.SECONDS);
// نرخ ثابت (fixed RATE): هر ۱۰ ثانیه، صرفنظر از مدت وظیفه
sched.scheduleAtFixedRate(this::poll, 0, 10, TimeUnit.SECONDS);
// تأخیر ثابت (fixed DELAY): ۱۰ ثانیه بعد از هر اتمام صبر میکند — بدون انباشت
sched.scheduleWithFixedDelay(this::poll, 0, 10, TimeUnit.SECONDS);
تفاوت این دو متد را باید در استخوانت حس کنی چون یک تمایز حیاتی است:
scheduleAtFixedRate مثل قطاری است که ساعتِ حرکتش ثابت است: هر ساعت رأس ساعت راه میافتد، مهم نیست سفرِ قبلی چقدر طول کشیده. اگر یک اجرا از دورهٔ ۱۰ ثانیهای بیشتر طول بکشد، اجرای بعدی بلافاصله پشتسرش شروع میشود تا عقبماندگی را جبران کند (اما هرگز روی یک وظیفهٔ واحد همزمان نمیشوند). در مقابل scheduleWithFixedDelay مثل اتوبوسی است که بعد از هر بار برگشتن به پایانه، ۱۰ ثانیه استراحت میکند و بعد دوباره راه میافتد — یعنی فاصله همیشه بین پایان یکی و شروع بعدی است، پس هرگز انباشت نمیشود.
scheduleAtFixedRate: یک دورهٔ ثابت را هدف میگیرد. اگر اجرا از دوره فراتر رود، اجراهای بعدی پشتسرهم برای جبران اجرا میشوند.scheduleWithFixedDelay: یک فاصلهٔ ثابت بین پایان یک اجرا و شروع بعدی میگذارد، پس اجراها هرگز روی هم نمیریزند.
این یکی واقعاً ساعتها از عمرت را میتواند بگیرد. اگر یک وظیفهٔ زمانبندیشده یک استثنای گرفتهنشده پرتاب کند، آن وظیفه لغو میشود و دیگر هرگز اجرا نمیشود — استخر آن را بیصدا سرکوب میکند و هیچ اجرای بعدیای رخ نمیدهد، بدون هیچ خطایی در لاگ. تصور کن یک کِران (cron) داری که ماهها کار کرده و یک روز «همینطوری» میایستد. علتش تقریباً همیشه همین است. همیشه بدنهٔ وظایف دورهای را در try/catch بپیچ:
sched.scheduleWithFixedDelay(() -> {
try { doWork(); }
catch (Throwable t) { log.error("poll failed, will retry next tick", t); }
}, 0, 10, TimeUnit.SECONDS);
خاموشسازی: shutdown() در برابر shutdownNow()
یک نکتهٔ مهم: یک ExecutorService را باید خودت خاموش کنی، وگرنه JVM شما اصلاً خارج نمیشود! چون نخهای کارگرِ استخر بهطور پیشفرض نخِ غیر-daemon (نخِ کاربر) هستند و تا زندهاند JVM را زنده نگه میدارند. (پایینتر daemon را توضیح میدهم.)
دو راه خاموشسازی داری:
| متد | وظیفهٔ جدید میپذیرد؟ | وظایف قبلاً-صفشده؟ | وظایف در حال اجرا؟ |
|---|---|---|---|
shutdown() |
نه | تا پایان اجرا میشوند | تمام میشوند |
shutdownNow() |
نه | بهصورت List<Runnable> برگردانده میشوند، اجرا نمیشوند |
interrupt میشوند |
shutdown() یعنی «درِ ورودی را قفل میکنیم، مشتری تازه نمیپذیریم، اما هر کس داخل است و سفارش داده، غذایش را کامل میگیرد». shutdownNow() یعنی «همین الان تعطیل! آشپزها دست از کار بکشید (interrupt)، و سفارشهایی که هنوز شروع نشده روی پیشخوان بماند» — آن سفارشهای شروعنشده را بهصورت یک لیست به تو برمیگرداند تا خودت تصمیم بگیری با آنها چه کنی.
توجه کن که shutdownNow() وظایفِ در حال اجرا را interrupt میکند (پرچم وقفهشان را ست میکند)، اما نمیتواند وظیفهای را که به وقفه بیاعتناست بهزور متوقف کند. اصطلاحِ متعارف و درستِ خاموشسازیِ باوقار این است:
void shutdownGracefully(ExecutorService pool) {
pool.shutdown(); // پذیرش را قطع کن، بگذار وظایف صف تمام شوند
try {
if (!pool.awaitTermination(30, TimeUnit.SECONDS)) {
pool.shutdownNow(); // بهزور: وظایف در حال اجرا را interrupt کن
if (!pool.awaitTermination(10, TimeUnit.SECONDS))
log.error("pool did not terminate");
}
} catch (InterruptedException e) {
pool.shutdownNow();
Thread.currentThread().interrupt(); // پرچم را بازیابی کن
}
}
دو نکتهٔ ظریف: اول، awaitTermination فقط یک مقدار بولی برمیگرداند و خودش خاموشسازی را فعال نمیکند — باید حتماً اول shutdown() یا shutdownNow() را صدا بزنی، وگرنه تا ابد منتظر میمانی. دوم، از جاوا ۱۹ به بعد ExecutorService اینترفیس AutoCloseable را پیادهسازی میکند؛ متد close() مثل «shutdown() سپس انتظار» رفتار میکند، پس میتوانی استخر را در یک try-with-resources بگذاری و خیالت راحت باشد که خودکار بسته میشود.
مدل وقفه (لغو تعاونی)
رسیدیم به پرسش سوم: یک نخ چطور متوقف میشود؟ و اینجا جاوا یک تصمیم عمدی و مهم گرفته: هیچ راه امنی برای کشتنِ اجباریِ یک نخ وجود ندارد.
شاید بپرسی «مگر متد Thread.stop() نداریم؟» چرا، ولی منسوخ (deprecated) است و هرگز نباید استفادهاش کنی. دلیلش خطرناک است: stop() نخ را وسط کار قطع میکند و قفلهایش را وسطِ جهش (mutation) — یعنی وسطِ تغییرِ نیمهکارهٔ یک شیء — رها میکند. نتیجه؟ اشیایی که در حالت خراب و ناسازگار جا میمانند. مثل اینکه جراح را وسط عمل بهزور از اتاق بیرون بکشی. برای همین جاوا این در را بست.
پس لغو در جاوا تعاونی (cooperative) است. یعنی تو فقط درخواست لغو میکنی (با interrupt())، و نخِ هدف خودش انتخاب میکند که متوجه شود و کارش را تمیز جمع کند.
interrupt() مثل این است که آرام روی شانهٔ یک همکار که هدفون دارد بزنی و بگویی «هر وقت تونستی، کارتو تموم کن و بیا». تو نمیتوانی مغزش را خاموش کنی؛ فقط یک سیگنال میفرستی. اگر او آدمِ همکاری باشد، سرِ اولین فرصت متوجه میشود و تمیز جمع میکند. اگر بیاعتنا باشد و اصلاً چک نکند، هیچ اتفاقی نمیافتد. لغوِ تعاونی یعنی همین: همکاریِ نخِ هدف لازم است.
وقفه در واقع فقط یک پرچمِ بولیِ بهازای هر نخ است، بهعلاوهٔ چند قاعده:
thread.interrupt()— پرچم را ست میکند (سیگنال را میفرستد).Thread.interrupted()— ایستا (static) است؛ پرچمِ نخِ فعلی را میخواند و پاک میکند.thread.isInterrupted()— پرچم را میخواند بدون اینکه پاکش کند.- متدهای مسدودکنندهای که در امضایشان
throws InterruptedExceptionدارند (sleep,wait,join,BlockingQueue.take,Future.get) وقتی interrupt شوند، پرچم را پاک میکنند و یکInterruptedExceptionپرتاب میکنند.
حالا وقتی یک InterruptedException میگیری، دقیقاً دو پاسخِ درست داری:
- آن را منتشر کن (بگذار به لایههای بالاتر بالا برود)، یا
- پرچم را بازیابی کن و از کار خارج شو، تا فراخوانهای بالای پشته هنوز بتوانند لغو را ببینند:
try {
queue.take();
} catch (InterruptedException e) {
Thread.currentThread().interrupt(); // هرگز بیصدا نبلع
return; // واحد کار را متوقف کن
}
بدترین کاری که میتوانی بکنی این است که InterruptedException را با یک catch خالی بگیری و هیچ کاری نکنی. یادت باشد: وقتی این استثنا پرتاب میشود، پرچم وقفه پاک شده. حالا اگر تو هم آن را ببلعی، تنها سیگنالی را که به دنیا میگفت «بایست» کاملاً نابود کردهای — و حالا وظیفهات دیگر قابللغو نیست و حتی shutdownNow هم نمیتواند متوقفش کند. این یکی از پرتکرارترین باگهایی است که در بازبینیِ کدِ سنیور علامتگذاری میشود. هرگز، هرگز بیصدا نبلع.
یک نکتهٔ آخر: متدهای مسدودکننده خودشان InterruptedException میدهند، اما اگر یک حلقهٔ CPU-محور داری که هیچ متد مسدودکنندهای صدا نمیزند، باید پرچم را دستی چک کنی:
while (!Thread.currentThread().isInterrupted()) {
doOneChunk();
}
// حلقه سرِ interrupt تمیز خارج میشود — لغو تعاونی
ThreadLocal — و چرا در استخرها نشت میکند
ThreadLocal<T> به هر نخ یک نسخهٔ خصوصیِ خودش از یک متغیر میدهد. این مکانیزم استانداردِ نگهداشتنِ «زمینهٔ هر درخواست» است — مثل شناسهٔ تراکنش، کاربرِ جاری، یا یک نمونهٔ SimpleDateFormat — بدون اینکه مجبور باشی آن را از میان همهٔ متدها دستبهدست کنی.
ThreadLocal مثل کمدِ شخصیِ هر کارمند در محل کار است. همه به «کمد» دسترسی دارند (متغیر مشترک است)، اما داخلِ کمدِ هر کس فقط مالِ خودش است و کسی محتوای کمدِ دیگری را نمیبیند. این عالی است تا وقتی که… کارمندها عوض نشوند. اگر همان کمد را بعداً به یک کارمند تازه بدهی بدون اینکه خالیاش کنی، وسایلِ نفرِ قبلی هنوز داخلش است. و دقیقاً همین اتفاق در استخرِ نخ میافتد.
private static final ThreadLocal<SimpleDateFormat> FMT =
ThreadLocal.withInitial(() -> new SimpleDateFormat("yyyy-MM-dd"));
مقادیرِ ThreadLocal در یک نقشه (map) داخلِ خودِ شیءِ Thread ذخیره میشوند که کلیدش همان شیءِ ThreadLocal است. این در نخهای استخرشده دو خطرِ جدی میسازد:
نخهای استخر عمرِ طولانی دارند و بازاستفاده میشوند. مقداری که در حین سرویسدادن به درخواستِ A ست کردی، هنوز همانجاست وقتی همان نخ بعداً به درخواستِ B سرویس میدهد. حالا درخواستِ B زمینهٔ درخواست A را میبیند — مثلاً هویتِ کاربرِ A! این هم باگ درستی است هم یک باگِ امنیتی جدی. قاعدهٔ آهنین: اگر در یک نخِ استخرشده یک ThreadLocal را set میکنی، باید آن را در یک بلوکِ finally با remove() پاک کنی:
try {
RequestContext.CURRENT.set(ctx);
handle(request);
} finally {
RequestContext.CURRENT.remove(); // در استخرها اجباری
}
نقشهٔ درونیِ ThreadLocal (یعنی ThreadLocalMap) برای کلیدهایش (خودِ شیءِ ThreadLocal) از ارجاعِ ضعیف (weak reference) استفاده میکند اما برای مقادیر از ارجاعِ قوی (strong reference). حالا فکر کن: اگر خودِ شیءِ ThreadLocal دیگر جایی ارجاع نداشته باشد اما نخ زنده بماند (که نخهای استخر دقیقاً همینطورند)، مقدار قابل جمعآوری توسط زبالهروب (garbage collector) نیست — تا وقتی آن خانهٔ نقشه دوباره استفاده شود یا نخ بمیرد. اگر اشیای بزرگی را در ThreadLocal کش کرده باشی و استخرت طولانیعمر باشد، هیپ آرامآرام پر میشود. remove() این را هم رفع میکند.
شاید فکر کنی InheritableThreadLocal راهحل است — این نوع، مقدار را هنگام ساختِ نخِ فرزند به آن کپی میکند. اما نخهای استخر برای هر وظیفه از نو ساخته نمیشوند (نکتهٔ کلِ استخر همین است!)، پس زمینه را به وظایفِ استخر منتقل نمیکند. برای انتقالِ درست زمینه در استخرها به بستهبندیِ (wrapping) صریح نیاز داری، یا به ScopedValue که در جاوا ۲۱ آمده.
نخهای daemon
هر نخ در جاوا یا یک نخِ کاربر (user thread) است یا یک نخِ daemon. قاعدهٔ کلیدی: JVM دقیقاً وقتی خارج میشود که آخرین نخِ غیر-daemon تمام شود — و در آن لحظه، هر نخِ daemonی که هنوز در حال کار باشد را بیرحمانه رها میکند (بدون هیچ تضمینی که بلوکهای finally اجرا شوند).
نخِ کاربر مثل کارمند رسمی است: تا او کاری دارد، ساختمان (JVM) باز میماند. نخِ daemon مثل نگهبانِ شب یا سیستمِ تهویه است: کارِ پسزمینه انجام میدهد، اما لحظهای که آخرین کارمندِ رسمی برود، چراغها خاموش میشود و نگهبان هم — چه بخواهد چه نخواهد، حتی اگر وسطِ کاری باشد — بیرون میماند. برای همین کارِ حیاتی را هرگز به نگهبانِ شب نمیسپاری.
Thread t = new Thread(task);
t.setDaemon(true); // باید پیش از start() ست شود، وگرنه IllegalThreadStateException
t.start();
- از نخِ daemon برای زیرساختِ پسزمینه استفاده کن که هرگز نباید برنامه را زنده نگه دارد: فلاشکردنِ متریکها، تایمرهای پاکسازیِ کش، نظرسنجیهای ضربان (heartbeat).
- هرگز از نخِ daemon برای کاری که باید تمام شود یا باید پاکسازی کند استفاده نکن (نوشتن یک فایل، فلاشکردن یک تراکنش) — چون JVM ممکن است وسطِ نوشتن آن را بکشد.
- یک نخ بهطور پیشفرض وضعیتِ daemonِ والدش را به ارث میبرد. برای همین است که نخهای
Executorsنخِ کاربر (غیر-daemon) هستند و تا وقتی استخر را خاموش نکنی JVM را زنده نگه میدارند. اگر نخهای daemon در استخر میخواهی، باید یکThreadFactoryبدهی کهsetDaemon(true)را صدا بزند.
نخهای مجازی (Java 21+) — جایی که ریاضیِ اندازهگذاری بخار میشود
و حالا بزرگترین تغییرِ سالهای اخیر. جاوا ۲۱ نخهای مجازی (virtual threads) را دائمی کرد. بگذار اول شهودش را بسازم.
تصور کن یک شرکتِ پیک داری. نخِ سکویی (platform thread) مثل یک ماشینِ گرانقیمتِ شرکت است: تعدادش محدود است و هر کدام گران. اگر پیکی برود دمِ در و منتظرِ باز شدنِ در بماند، آن ماشینِ گران همانجا بیکار پارک است. نخِ مجازی مثل یک پیکِ کمهزینه است: وقتی به دری میرسد و باید منتظر بماند، از ماشین پیاده میشود (unmount) و ماشین را آزاد میکند تا پیکِ دیگری سوارش شود؛ وقتی در باز شد، دوباره سوارِ یک ماشینِ آزاد میشود. تعداد ماشینها کم است (نخهای حاملِ سکویی = carrier)، اما تعداد پیکها میتواند میلیونها باشد.
پس نخِ مجازی نخی است که خودِ JVM آن را روی یک استخرِ کوچک از نخهای حاملِ (carrier) سکویی زمانبندی میکند؛ وقتی روی I/O مسدود میشود، بهجای اینکه یک نخِ سیستمعامل را مسدود کند، پیاده میشود (unmount) — در فضای کاربر پارک میشود و حاملش را آزاد میکند. آنقدر ارزاناند که میتوانی میلیونها بسازی.
// هر وظیفه نخ مجازی خودش را میگیرد؛ این یک استخر نیست که اندازهاش کنی
try (var exec = Executors.newVirtualThreadPerTaskExecutor()) {
for (var req : requests)
exec.submit(() -> handleBlocking(req)); // فراخوانهای مسدودکننده اینجا اشکالی ندارند
} // close() منتظر همهٔ وظایف میماند (AutoCloseable)
این جانِ کلام است. برای کارِ I/O-محور، دیگر استخر را اندازهگذاری نمیکنی. کدِ سرراستِ مسدودکننده مینویسی، یک نخِ مجازی بهازای هر وظیفه، و میگذاری JVM آنها را روی حاملهای کم مالتیپلکس (چندگانهسازی) کند. یادت هست فرمول گوتز؟ آن فرمول در واقع یک راهِ دور زدنِ مشکلِ «نخِ سکوییِ گران» بود. نخِ مجازی همان محدودیتی را که فرمول برایش جبران میکرد، از ریشه حذف میکند.
اما چند مورد که هنوز میگزند:
اگر یک نخِ مجازی داخلِ یک بلوکِ synchronized (یا وسطِ یک فراخوانِ بومی/JNI) باشد و همانجا مسدود شود، دیگر نمیتواند از حاملش پیاده شود — به آن حامل پین (pin) میشود، یعنی به آن میخکوب میماند. اگر بهقدر کافی نخ به حاملها پین شوند، استخرِ حاملها گرسنه میشود و توان عملیاتی میخوابد. راهحل: حولِ I/Oِ مسدودکننده بهجای synchronized از ReentrantLock استفاده کن که به نخِ مجازی اجازهٔ پیادهشدن میدهد. (نسخههای بعدیِ JDK، یعنی ۲۴ به بعد، بیشترِ پینشدنِ ناشی از synchronized را حذف کردهاند، اما روی خودِ جاوا ۲۱ این مشکل واقعی است.)
- استخرشان نکن. ساختنِ
newFixedThreadPoolاز نخهای مجازی کلِ فلسفه را نقض میکند؛ همیشه بهازای هر وظیفه یکی بساز. - برای کارِ CPU-محور هیچ سودی ندارند. نخِ مجازی به انتظار کمک میکند نه به محاسبه. برای محاسبهٔ محض همان استخرِ نخِ سکوییِ اندازهگرفتهشده به هستهها را نگه دار.
تلّهها و نکتههای رایج (رگبار سریع)
بیا همهٔ تلّههایی را که در فصل دیدیم یکجا مرور کنیم — اینها همانهاییاند که در تولید خوندل درست میکنند:
- صدا زدنِ
run()بهجایstart()— بدون نخ جدید. - صفِ بیکرانِ
newFixedThreadPool→ OOM زیر بار. maximumPoolSizeنادیده گرفته میشود چون صف بیکران است.- بلعیدنِ
InterruptedException→ وظایفِ غیرقابللغو. - فراموشکردنِ
ThreadLocal.remove()در استخر → زمینهٔ نشتکرده/کهنه. - نامگذارینکردنِ نخها → thread dumpهای بیفایده در تولید. همیشه
ThreadFactoryبده. - استثناهای وظایفِ
submitشده تاget()بیصدا گم میشوند. - وظیفهٔ زمانبندیشده یکبار پرتاب کند → برای همیشه بیصدا میایستد.
- اشتراکِ شیءِ غیرِ thread-safe (
SimpleDateFormat،Randomزیر رقابت) بین نخهای استخر.
بهترین روشها
۱. ThreadPoolExecutor را صریح با صف کراندار + سیاست ردّ انتخابی بساز؛ در تولید از کارخانههای بیکرانِ Executors بپرهیز.
۲. نخهایت را نامگذاری کن از طریق ThreadFactory — خودِ آیندهات هنگام خواندنِ thread dump سپاسگزار میشود.
۳. استخرهای CPU را به cores اندازه بگیر؛ استخرهای I/O را با فرمولِ انتظار/محاسبه، سپس در گلوگاهِ واقعیِ پاییندستی سقف بگذار.
۴. همیشه هنگام پایینآمدن shutdown() + awaitTermination() + shutdownNow().
۵. با InterruptedException مثل سیگنالی درجهیک برخورد کن: منتشر کن یا بازیابی-و-خروج، هرگز نبلع.
۶. در استخرها هر ThreadLocal را در finally با remove() پاک کن.
۷. CompletableFuture/همزمانیِ ساختاریافته را بر زنجیرههای get()ِ مسدودکننده ترجیح بده.
۸. در جاوا ۲۱ به بعد، برای fan-outِ I/O-محور از نخهای مجازی استفاده کن و تنظیمِ دستیِ اندازهٔ استخر را کنار بگذار.
پرسشهای مصاحبه
حالا بیا هر چه یاد گرفتی را در قالبِ پرسشهای واقعیِ مصاحبه تمرین کنیم. هر کدام را اول خودت جواب بده، بعد پاسخ را بخوان.
BLOCKED یعنی نخ مشخصاً بر سرِ قفلِ مانیتورِ synchronized رقابت میکند. WAITING/TIMED_WAITING یعنی park شده و منتظرِ سیگنال است. RUNNABLE یعنی JVM آن را قادر به اجرا میداند. JVM انسدادِ سطحِ سیستمعامل روی I/O را نمیبیند، پس نخی که در خواندنِ مسدودکنندهٔ سوکت گیر کرده هنوز RUNNABLE برچسب میخورد — انسداد زیرِ دیدِ JVM رخ میدهد.
start() دوم IllegalThreadStateException پرتاب میکند — نخ یکبارمصرف است. صدا زدنِ run() بدنه را همریسمانی روی نخِ فعلی اجرا میکند؛ نخِ جدیدی ساخته نمیشود. این باگی تصادفیِ رایج است.
فقط ۴. با صفِ بیکران، وظایف همیشه با موفقیت صف میشوند، پس استخر هرگز به شرطِ «صف پر» که ساختِ نخِ بالای core را فعال میکند نمیرسد. maximumPoolSize عملاً مرده است. برای استفاده از max به صفِ کراندار نیاز داری.
(۱) اگر نخها < core، نخ جدید بساز. (۲) وگرنه سعی کن صف کنی. (۳) اگر صفکردن شکست خورد (صف پر) و نخها < max، نخِ غیر-core بساز. (۴) وگرنه handlerِ ردّ را فراخوان کن. نکتهٔ ظریف: صفکردن پیش از رشد فراتر از core امتحان میشود، به همین دلیل صفِ بیکران شما را در core سقف میکند.
AbortPolicy پرتاب میکند (fail-fast، پیشفرض). DiscardPolicy بیصدا میاندازد. DiscardOldestPolicy سرِ صف را میاندازد و دوباره امتحان میکند. CallerRunsPolicy وظیفه را روی نخِ ثبتکننده اجرا میکند — این فشار برگشتی میدهد چون ثبتکننده مشغول است، نمیتواند بیشتر ثبت کند و ورودی را throttle میکند.
threads = cores × util × (1 + wait/compute) = 16 × 1.0 × (1 + 190/10) = 16 × 20 = 320. اما در سقفِ محدودیتِ همزمانیِ واقعیِ پاییندستی هم کران بگذار — ۳۲۰ نخ که به سرویسی با ۵۰ همزمانِ مجاز میکوبند فقط صف میکشند؛ به اندازهٔ گلوگاه بگیر.
پرچمِ وقفه هنگامِ پرتابِ استثنا پاک میشود. اگر آن را بگیری و کاری نکنی، تنها سیگنالِ لغو را نابود کردهای — وظیفه غیرقابلinterrupt میشود و shutdownNow نمیتواند متوقفش کند. درست: یا دوباره پرتاب/منتشرش کن، یا Thread.currentThread().interrupt() را صدا بزن تا پرچم بازیابی شود و سپس از کار خارج شو.
shutdown(): پذیرشِ وظیفهٔ جدید را قطع میکند، میگذارد وظایفِ صفشده و در حال اجرا تمام شوند. shutdownNow(): پذیرش را قطع میکند، وظایفِ صفشدهٔ شروعنشده را بهصورت List<Runnable> برمیگرداند (هرگز اجرا نمیشوند) و وظایفِ در حال اجرا را interrupt میکند (که فقط اگر به وقفه احترام بگذارند متوقفشان میکند).
درستی: نخهای استخر بازاستفاده میشوند، پس مقداری که برای یک وظیفه ست شده به وظیفهٔ بعدی روی همان نخ نشت میکند مگر remove() کنی — زمینهٔ یک درخواست (مثلاً هویتِ کاربر) به دیگری نشت میکند. حافظه: کلیدهای ThreadLocalMap ضعیفاند اما مقادیر قوی؛ نخِ طولانیعمرِ استخر مقدار را نگه میدارد حتی بعد از رفتنِ ThreadLocal، تا خانه بازاستفاده شود. remove() در finally هر دو را رفع میکند.
ExecutorService pool = Executors.newSingleThreadExecutor();
Future<?> f = pool.submit(() -> { throw new RuntimeException("boom"); });
System.out.println("submitted");
// بدون فراخوان f.get()
pool.shutdown();
submitted را چاپ میکند و هیچچیز دربارهٔ «boom». استثناهای وظایفِ submitشده در Future گرفته میشوند و فقط روی f.get() رو میشوند (بهصورت ExecutionException). بدون get()، استثنا بیصدا بلعیده میشود. (اگر از execute() استفاده کرده بودی، به UncaughtExceptionHandler میرسید و stack trace چاپ میکرد.)
scheduleAtFixedRate یک دورهٔ ثابت را هدف میگیرد؛ اگر اجرا از دوره فراتر رود، اجرای بعدی بلافاصله بعدش شروع میشود (برای جبران burst میکند، هرگز روی وظیفهٔ واحد همپوشانی نمیکند). scheduleWithFixedDelay همیشه تأخیرِ کامل را بعد از اتمام میگذارد، پس اجراها هرگز انباشته نمیشوند. برای نظرسنجیای که نباید همپوشانی یا هجوم کند، از تأخیرِ ثابت استفاده کن.
استثنای گرفتهنشده پرتاب کرد. وظیفهٔ زمانبندیشدهای که پرتاب کند بیصدا لغو میشود و دیگر هرگز زمانبندی نمیشود؛ استثنا در Futureِ (نادیدهگرفتهشده) دام میافتد. بدنهٔ وظیفه را در try/catch بپیچ تا زنده بماند.
وضعیتِ daemon وقتی نخ زنده شد تغییرناپذیر است — ستکردنش بعد از start() استثنای IllegalThreadStateException میدهد. یک daemon وقتی داده از دست میدهد که تنها اجراکنندهٔ عملیاتی حیاتی باشد و همهٔ نخهای کاربر خارج شوند: JVM خاتمه مییابد و daemon را وسطِ عملیات رها میکند، بدون تضمینِ اجرای finally/پاکسازی. هرگز کارِ باید-تکمیلشود را روی daemonها نگذار.
پینشدن (pinning). I/Oِ مسدودکننده درونِ بلوکِ synchronized (یا فراخوانِ بومی) نشسته، پس نخِ مجازی هنگامِ مسدودشدن نمیتواند از حاملش پیاده شود. با حاملهای کم، نخهای مجازیِ پینشده استخر را گرسنه میکنند. رفع: synchronized حولِ فراخوانِ مسدودکننده را با ReentrantLock جایگزین کن که به نخِ مجازی اجازهٔ پیادهشدن میدهد. (همچنین بررسی کن نخهای مجازیِ بهازای-وظیفه استفاده میکنی، نه استخرشده.)
برای کارِ CPU-محور — موازیسازیای فراتر از تعدادِ حاملها که به هستهها کران دارد اضافه نمیکنند. وظایفِ محاسبهسنگین باید از استخرِ نخِ سکوییِ اندازهگرفتهشده به هستهها استفاده کنند. نخهای مجازی فقط وقتی میبرند که نخها وقتشان را منتظر بگذرانند.
نکاتِ سنیور و موارد پیشرفته
تا اینجا مدل ذهنی نخ، Executor، اندازهگیری pool، لغو تعاونی و virtual threads را داری. حالا میرویم سراغ لایهای که فقط بعد از چند شبِ on-call و چند post-mortem واقعی یاد میگیری: چیزهایی که در کد درست به نظر میرسند ولی زیر بار prod میترکند. این بخش هیچکدام از مطالب قبلی را تکرار نمیکند؛ فقط شکافها را پر میکند.
۱. مدل حافظه (JMM): چه garanteeای مجانی از Executor میگیری و چرا برای خواندن نتیجهٔ task نیازی به volatile نداری.
۲. availableProcessors() در container دروغ میگوید — بزرگترین اشتباه sizing در Kubernetes.
۳. ForkJoinPool.commonPool() و parallel stream یک pool مشترکاند و blocking داخلشان کل JVM را قحطیزده میکند.
۴. CompletableFuture درست — کدام thread callback را اجرا میکند، join در برابر get، جمعکردن نتایج.
۵. Thread-starvation deadlock (pool تو در pool).
۶. Observability و hookها، propagation کانتکست/MDC، انواع queue.
۷. دامهای @Async در Spring.
۸. پلیبوک واقعی virtual threads: Semaphore جای sizing، pinning در JDK جدید، ScopedValue و structured concurrency.
۹. هفت سؤال سختِ سنیور با پاسخ کامل.
مدل حافظه: happens-beforeای که مجانی میگیری
یک سؤال ظریف: task روی نخ B مقداری را حساب میکند و در یک فیلد معمولی (بدون volatile) مینویسد؛ نخ A بعد از future.get() آن فیلد را میخواند. آیا A حتماً مقدار تازه را میبیند یا ممکن است مقدار کهنه/نیمهساخته ببیند؟
مستندات پکیج java.util.concurrent دو رابطهٔ happens-before را صریحاً تضمین میکنند:
۱. هر چیزی که قبل از submit/execute روی نخ فرستنده انجام دادی، happens-before شروع اجرای task است. یعنی task همهٔ نوشتههای قبل از submit را میبیند.
۲. هر چیزی که task انجام میدهد، happens-before بازگشتِ Future.get() است. یعنی بعد از get() نخ A تضمیناً همهٔ نوشتههای task را بهروز میبیند.
پس نیازی به volatile یا synchronized برای انتقال نتیجهٔ task نداری؛ خودِ مرز submit→run→get این visibility را میسازد. همین ضمانت برای BlockingQueue (put/take)، CountDownLatch، و اتمام یک نخ نسبت به join() هم برقرار است.
اگر task را با execute() بفرستی و نتیجه را بدون هیچ سازوکار همگامسازی (نه get، نه queue، نه latch، نه فیلد volatile) از نخ دیگری بخوانی، هیچ رابطهٔ happens-before نداری و ممکن است برای همیشه مقدار کهنه یا حتی object نیمهساخته ببینی. باگش هم غیرقطعی است و در تست دیده نمیشود.
availableProcessors() در container دروغ میگوید
فرمولهای sizing فصل قبل همه به «تعداد coreها» تکیه دارند. اما آن عدد از کجا میآید؟ از Runtime.getRuntime().availableProcessors(). مشکل اینجاست که در Kubernetes این عدد اغلب همان چیزی که فکر میکنی نیست.
node تو ۳۲ core دارد، ولی به pod یک cpu limit مثلاً 500m (نصف یک core) دادهای. JVM از JDK 10 به بعد container-aware است و cgroup را میخواند؛ با quota زیر یک core، availableProcessors() مقدار ۱ برمیگرداند. حالا سه فاجعه:
- اگر pool را با عددِ ثابت ۳۲ ساخته بودی، ۳۲ نخ روی سهمیهٔ یک core دائم context-switch میکنند و throttle میشوی.
ForkJoinPool.commonPool()که parallelismاشavailableProcessors()-1است، عملاً ۰ (یعنی یک نخ) میشود و همهٔ parallel streamها سریال اجرا میشوند.- برعکس، اگر limit نگذاری ولی JVM قدیمی/غلط پیکربندی باشد، ممکن است ۳۲ ببیند و روی همهٔ همسایهها اثر بگذارد.
- cgroups v2 از JDK 17 به بعد کامل پشتیبانی میشود؛ روی JDKهای قدیمی و k8s جدید حواست باشد.
- اگر limit را کسری گذاشتی ولی میخواهی JVM عدد مشخصی ببیند،
-XX:ActiveProcessorCount=Nرا صریح ست کن. - عدد واقعی را در startup لاگ کن:
log.info("cores={}", Runtime.getRuntime().availableProcessors()). همین یک خط ساعتها debugging را حذف میکند. - ترجیحاً
cpu request == cpu limitبگذار تا رفتار قابلپیشبینی شود.
commonPool() و parallel streamها یک pool مشترکاند
stream().parallel() و CompletableFuture.supplyAsync(task) (نسخهٔ بدون executor) هر دو روی یک pool سراسری اجرا میشوند: ForkJoinPool.commonPool(). این pool کوچک است (به اندازهٔ coreها) و برای کار CPU-bound کوتاه ساخته شده.
اگر داخل یک parallelStream().map(...) یک فراخوانی blocking (HTTP، DB، Thread.sleep) بگذاری، آن worker نخِ commonPool را قفل میکند. چون این pool در کل JVM مشترک است، حالا هر بخش دیگری از برنامه که به parallel stream یا supplyAsync تکیه دارد گرسنه میماند و درخواستهای بیربط کند میشوند. کلاسیکترین حادثه: یک endpoint گزارشگیری با parallel stream، کل سرویس را با هم میخواباند.
دو راهحل درست:
// راه ۱: parallel stream را داخل یک ForkJoinPool اختصاصی اجرا کن
ForkJoinPool custom = new ForkJoinPool(16);
List<R> out = custom.submit(() ->
items.parallelStream().map(this::blockingCall).toList()
).get();
// راه ۲ (برای CompletableFuture): همیشه executor خودت را پاس بده
CompletableFuture.supplyAsync(this::blockingCall, myIoPool);
اگر مجبوری داخل FJP block کنی، بهجای block ساده از ForkJoinPool.managedBlock(...) استفاده کن. این interface به pool خبر میدهد که «من دارم block میشوم»، و pool موقتاً یک worker جبرانی راه میاندازد تا parallelism حفظ شود. CompletableFuture هم داخلش از همین ManagedBlocker برای join() روی worker نخهای FJP استفاده میکند.
CompletableFuture درست: کجا callback اجرا میشود
فصل قبل CompletableFuture را در حد «non-blocking و composable» معرفی کرد. اما نکتهٔ سنیوری این است: callback کجا اجرا میشود؟
- نسخهٔ بدون
Async(مثلthenApply): callback روی نخی اجرا میشود که stage قبلی را کامل کرده — یا حتی روی نخِ فراخواننده اگر آن لحظه از قبل complete شده باشد. غیرقابلپیشبینی و خطرناک برای کارِ سنگین. - نسخهٔ
Asyncبدون executor (thenApplyAsync): رویcommonPool()— با همان دام قحطی بالا. - نسخهٔ
Asyncبا executor: روی pool خودت. این حالت پیشفرضِ درستِ prod است.
CompletableFuture
.supplyAsync(() -> fetchUser(id), ioPool) // I/O روی pool ما
.thenApplyAsync(this::enrich, cpuPool) // CPU روی pool جدا
.exceptionally(ex -> User.anonymous()) // fallback
.thenAccept(this::send);
اگر تابعت خودش یک CompletableFuture برمیگرداند، thenApply به CompletableFuture<CompletableFuture<T>> تودرتو میرسد؛ برای flatten از thenCompose (معادل flatMap) استفاده کن. get() استثنای checked میدهد؛ join() نسخهٔ unchecked است (CompletionException) و داخل زنجیره تمیزتر. برای جمعکردن چند future: allOf(a,b,c).thenApply(...) — ولی allOf مقدار Void برمیگرداند، پس بعدش تکتک join بزن. برای خطا: exceptionally فقط مسیر خطا، handle هر دو مسیر، و whenComplete یک finally است که مشاهده میکند ولی تغییر نمیدهد.
Thread-starvation deadlock: pool تو در pool
این یکی از موذیترین deadlockهاست و در تست هیچوقت دیده نمیشود چون فقط زیر بار رخ میدهد.
توالی: pool تو کراندار است (مثلاً ۱۰ نخ). یک task روی این pool، خودش subtaskهایی را به همان pool submit میکند و بعد get() میزند و منتظر نتیجهٔ آنها میماند. اگر ۱۰ taskِ والد همزمان همهٔ ۱۰ نخ را بگیرند و هرکدام منتظر subtask خودش باشد، دیگر هیچ نخی برای اجرای subtaskها آزاد نیست — همه منتظرِ همه؛ قفل کامل.
نمودار — چرا pool تودرتوی کراندار قفل میشود:
sequenceDiagram
participant P as Parent tasks (fill all N threads)
participant Q as Work queue
participant C as Child tasks
P->>Q: submit child, then block on get()
Q-->>C: waiting for a free thread
Note over P,C: all N threads held by parents<br/>no thread free to run children
Note over P,C: parents wait children, children wait threads = deadlock
یک taskی که روی یک pool اجرا میشود نباید به همان pool کار submit کند و روی نتیجهاش block شود. یا poolهای جدا برای لایههای مختلف بگذار، یا کل زنجیره را non-blocking با thenCompose بنویس، یا از virtual threads استفاده کن (که نخ کمیاب نیست و این کلاس deadlock را حذف میکند).
Observability و hookها
poolی که نمیتوانی ببینی یک روز غافلگیرت میکند. getActiveCount()، getQueue().size() و getCompletedTaskCount() را به Micrometer بده — صفِ همیشهپر یعنی pool کوچک یا downstream کند. و beforeExecute/afterExecute را در ThreadPoolExecutor override کن: جای عالی برای ست/پاککردن MDC، اندازهگیری زمان اجرا، و گرفتنِ استثناهایی که در submit گم میشوند (afterExecute هم Runnable و هم Throwable را میگیرد). دو دستهٔ دیگر: allowCoreThreadTimeOut(true) اجازه میدهد نخهای coreِ بیکار هم جمع شوند (poolهای کممصرف)، و prestartAllCoreThreads() نخهای coreِ lazy را از قبل warm میکند وقتی latencyِ اولین درخواستها مهم است.
ThreadLocalهای SLF4J MDC، SecurityContextHolder و trace-idهای tracing خودکار به نخ pool منتقل نمیشوند؛ نخ کارگر آنها را ندارد. راهحل: یک TaskDecorator (در Spring) یا wrapper که قبل از اجرا کانتکست را از نخ فرستنده کپی و بعدش پاک کند. با virtual threads این درد کمتر میشود ولی حذف نمیشود.
انتخاب نوع queue هم یک تصمیم مهندسی است: ArrayBlockingQueue پیشفرضِ امن prod است، PriorityBlockingQueue برای صف با اولویت، و DelayQueue برای زمانبندی — همه بهجای LinkedBlockingQueueی بیکران و خطرناک.
دامهای @Async در Spring
@Async قشنگترین «برو async» دنیاست تا وقتی یکی از اینها بزندت:
۱. self-invocation: @Async با یک proxy کار میکند. اگر یک متد public از همان bean متد @Async را صدا بزند، فراخوانی از proxy رد نمیشود و کاملاً سنکرون اجرا میشود — بدون هیچ خطا. باید از یک bean دیگر صدایش بزنی.
۲. گمشدن استثنا: اگر متد @Async بازگشت void داشته باشد و خطا بدهد، استثنا هیچجا نمیرود (مثل execute بدون Future) و در لاگ گم میشود؛ برای گرفتنش باید AsyncUncaughtExceptionHandler تعریف کنی، یا بازگشت را CompletableFuture کنی تا خطا در get دربیاید.
۳. executor پیشفرض: در Spring خالص (بدون Boot) اگر executor مشخص نکنی، fallback به SimpleAsyncTaskExecutor است که pool ندارد و برای هر فراخوانی یک نخ نو میسازد — زیر بار همان انفجار نخ. در Spring Boot خوشبختانه یک ThreadPoolTaskExecutor خودکار میسازد، ولی به آن هم default اعتماد نکن و صریح پیکربندی کن.
از Spring Boot 3.2 به بعد با spring.threads.virtual.enabled=true (روی Java 21+) هم executor @Async و هم نخهای Tomcat به virtual threads سوئیچ میشوند و کدِ blocking سنتیات بدون بازنویسی مقیاس میگیرد. ولی همان هشدارهای MDC و pinning همچنان برقرارند.
پلیبوک واقعی virtual threads
فصل قبل گفت با virtual threads دیگر pool را sizing نمیکنی. سؤال بعدیِ سنیور: پس جلوی downstream محدود را چطور بگیرم؟
وقتی یک virtual thread برای هر task میسازی، دیگر pool مرزی ندارد؛ ولی DB تو شاید فقط ۲۰ connection بدهد. اگر ۱۰٬۰۰۰ virtual thread همزمان به آن حمله کنند، connection pool منفجر یا timeout میشود. راهحل: concurrency را نه با اندازهٔ pool، بلکه با یک Semaphore به منبعِ کمیاب bound کن:
Semaphore db = new Semaphore(20);
db.acquire();
try { return jdbc.query(...); }
finally { db.release(); }
در دنیای virtual threads، Semaphore همان نقشی را دارد که قبلاً maximumPoolSize داشت: محافظ ظرفیت downstream.
با platform threadها یک ترفند رایج، cacheِ یک شیء گران (مثلاً بافر بزرگ) در ThreadLocal بود؛ چون نخها کم بودند. با میلیونها virtual thread این آنتیپترن است: هر نخ یک نسخه یعنی مصرف حافظهٔ انفجاری. با virtual threads یا شیء را per-task بساز، یا آن را pool کن، یا ScopedValue را در نظر بگیر.
در فصل گفتیم synchronized باعث pinning میشود. JEP 491 در JDK 24 دقیقاً این را رفع کرد: بلاک synchronized دیگر معمولاً carrier را گروگان نمیگیرد. فقط native/JNI و چند گوشهٔ نادر ماندهاند. برای پایش pinning، بهجای فلگ قدیمیِ -Djdk.tracePinnedThreads (که در JDK 24 deprecate شد) از رویداد JFR بهنام jdk.VirtualThreadPinned استفاده کن. تعداد carrierها را هم با -Djdk.virtualThreadScheduler.parallelism تنظیم میکنی.
ساخت دستی با Thread.ofVirtual().name("job-", 0).start(runnable) یا Thread.startVirtualThread(r). دو همراهِ مدرن الگوی fan-out را کامل میکنند: structured concurrency (در JDK 25 هنوز preview، JEP 505، حالا با StructuredTaskScope.open() و Joiner) چند subtask را در یک scope اجرا میکند تا همه با هم fail/cancel شوند و leak نخ نداشته باشی؛ و ScopedValue (در JDK 25 نهایی شد، JEP 506) جایگزینِ immutable و ارزانترِ ThreadLocal برای انتقال کانتکست به subtaskهاست — بدون نیاز به remove()، چون دامنهاش خودکار بسته میشود.
سؤالهای سختِ سنیور
نه. JMM تضمین میکند هرچه task انجام میدهد happens-before بازگشت Future.get() است، و هرچه قبل از submit نوشتی happens-before شروع task. پس مرز submit→run→get خودش visibility کامل میسازد و نیازی به volatile/synchronized برای انتقال نتیجه نیست. اما اگر بدون هیچ نقطهٔ همگامسازی (نه get، نه queue، نه latch) از fire-and-forget نتیجه بخوانی، هیچ ضمانتی نداری و مقدار کهنه محتمل است.
چون JVM از JDK 10+ container-aware است و cgroup را میخواند: با limit یک core، availableProcessors() مقدار ۱ (نه ۳۲) برمیگرداند. پس pool تو ۲ نخ میشود، و مهمتر، commonPool() که parallelismاش cores-1 است عملاً یکنخی میشود و parallel streamها سریال اجرا میشوند. راهحل: عدد واقعی را لاگ کن، در صورت نیاز -XX:ActiveProcessorCount را صریح ست کن، و cpu request == limit بگذار.
چون parallel stream روی ForkJoinPool.commonPool() سراسری اجرا میشود و فراخوانی blockingِ remote، نخهای این pool را قفل میکند؛ چون این pool در کل JVM مشترک است، بقیهٔ استفادهکنندهها گرسنه میمانند. راهحل ۱: stream را داخل یک ForkJoinPool اختصاصی submit کن. راهحل ۲ (بهتر): از parallel stream برای کار blocking استفاده نکن؛ بهجایش CompletableFuture.supplyAsync(task, dedicatedIoPool) یا virtual threads. اگر مجبوری داخل FJP block کنی، ManagedBlocker بزن تا pool worker جبرانی بسازد.
Thread-starvation deadlock. اگر ۱۰ taskِ والد همزمان هر ۱۰ نخ را بگیرند و هرکدام منتظر subtaskِ خودش باشد، هیچ نخی برای اجرای subtaskها نمیماند؛ والدها منتظر فرزندان، فرزندان منتظر نخ آزاد. راهحل: poolهای جدا برای لایههای مختلف، یا زنجیرهٔ non-blocking با thenCompose، یا virtual threads که نخ در آنها کمیاب نیست.
thenApply: callback روی همان نخی که stage قبلی را complete کرده (یا حتی نخ فراخواننده اگر از قبل complete بوده) — برای کار سنگین غیرقابلپیشبینی. thenApplyAsync بدون executor: روی commonPool (دام قحطی). thenApplyAsync با executor: روی pool تو — حالت درست prod. get() استثنای checked میدهد (ExecutionException)؛ join() نسخهٔ unchecked است (CompletionException) و داخل زنجیره تمیزتر.
چون @Async با proxy کار میکند؛ فراخوانی داخلی (self-invocation) از proxy رد نمیشود پس asyncسازی اتفاق نمیافتد و کد سنکرون میشود — باید از bean دیگری صدایش بزنی یا proxy را self-inject کنی. دربارهٔ خطا: متد @Async با بازگشت void مثل execute بدون Future است؛ استثنا هیچجا catch نمیشود و در لاگ گم میشود. راهحل: AsyncUncaughtExceptionHandler تعریف کن یا بازگشت را CompletableFuture کن تا خطا در join/get دربیاید.
concurrency را با یک Semaphore(20) دورِ فراخوانی downstream bound میکنی؛ Semaphore در دنیای virtual thread همان نقش maximumPoolSize را دارد. برای pinning: در JDK 24 با JEP 491 دیگر synchronized معمولاً pin نمیکند؛ فقط native/JNI مانده. برای پایش از رویداد JFR بهنام jdk.VirtualThreadPinned استفاده کن (فلگ قدیمی -Djdk.tracePinnedThreads در JDK 24 deprecate شد).
- مرز
submit → run → Future.getخودش happens-before میسازد؛ برای انتقال نتیجه به volatile نیازی نیست، ولی fire-and-forget بدون نقطهٔ همگامسازی امن نیست. - در container،
availableProcessors()از cgroup میآید و ممکن است بسیار کوچکتر از coreهای node باشد — همهٔ sizingها را میشکند. commonPool()و parallel streamها سراسریاند؛ blocking داخلشان کل JVM را قحطیزده میکند (executor اختصاصی یا ManagedBlocker). pool تودرتوی کراندار هم thread-starvation deadlock است.- در
CompletableFutureنسخهٔAsyncبا executor خودت را بده؛ pool را observable کن و MDC را با decorator منتقل کن. - دامهای
@Async: self-invocation سنکرون میشود،voidخطا را میبلعد، default ممکن است pool نداشته باشد. - با virtual threads: Semaphore جای sizing، ThreadLocalِ گران ممنوع، pinning در JDK 24 تقریباً حل شد؛ ScopedValue و structured concurrency الگوی مدرن fan-outاند.
- نخِ سکویی گران است (پشتهٔ ~۱MB + فراخوان سیستمی)؛ برای همین Executorها ثبتِ وظیفه را از ساختِ نخ جدا میکنند.
- شش حالتِ
Thread.Stateرا حفظ کن.BLOCKEDفقط رقابت بر سرِ مانیتور است؛ I/O در سطح OS بهصورتRUNNABLEدیده میشود. run()نخ نمیسازد،start()میسازد.Runnableنتیجه ندارد،Callableدارد و نتیجهاش را باFuture.get()میگیری — که مسدودکننده است و استثنا را تا لحظهٔget()قایم میکند.- در تولید
ThreadPoolExecutorرا دستی با صفِ کراندار و سیاستِ ردّ بساز. با صفِ بیکران،maximumPoolSizeبیاثر میشود. - CPU-محور را به تعداد هسته اندازه بگیر؛ I/O-محور را با
cores × (1 + wait/compute)، و در گلوگاهِ واقعی سقف بگذار. - برای زمانبندی، فرقِ fixed-rate و fixed-delay را بدان و بدنهٔ وظیفه را در try/catch بپیچ (وگرنه مرگِ خاموش).
- همیشه
shutdown()→awaitTermination()→shutdownNow(). لغو تعاونی است؛InterruptedExceptionرا هرگز نبلع. - در استخرها هر
ThreadLocalرا درfinallyباremove()پاک کن (درستی + امنیت + حافظه). - در جاوا ۲۱+ برای کارِ I/O-محور نخهای مجازی بهازای هر وظیفه بساز و اندازهگذاری را فراموش کن — فقط مراقبِ pinning باش و آنها را استخر نکن.
منابع: ThreadPoolExecutor (JDK 17)، Managing Throughput with Virtual Threads — inside.java، RejectedExecutionHandler — Baeldung
Come sit next to me. We are about to talk about one of those topics everyone dreads in interviews and bleeds over in production: concurrency. But I promise you — if you truly understand three things, this topic stops being scary magic and becomes something you hold in the palm of your hand. This whole chapter is those three things, built from the ground up, one small brick at a time.
First, a mental picture: a thread is like a worker who can follow one task. One worker means tasks happen one after another. Several workers means tasks can go forward at the same time. The entire argument of this chapter is about how you build those workers, how you hand them work, and how you politely say "that's enough, go home."
Here is the path we'll walk:
- What a thread actually is and why creating one is expensive (that expense is the whole reason Executors exist).
- The thread lifecycle — the six states a thread can be in.
- The unit of work:
Runnable,Callable, and the result-handleFuture. - The Executor framework and its heart,
ThreadPoolExecutor, with its seven parameters. - Pool-sizing math — where senior candidates earn their stripes.
- Scheduling (
ScheduledExecutorService) and graceful shutdown. - Cooperative cancellation (
InterruptedException) andThreadLocal, which leaks in pools. - Daemon threads, and finally virtual threads (Java 21+), which make the sizing math evaporate. At the end: fifteen interview questions with full answers. Ready? Let's go.
Part 0 — words you must know first
Before we jump into code, let's plant a few words we'll keep reusing, each with a small analogy. Keep these in the back of your mind and the rest of the chapter is easy.
Picture a restaurant. A cook is a thread: one person who can carry one order from start to finish. The maître d' who decides which cook works right now is the OS scheduler — because the kitchen only has a few stoves (CPUs/cores), not everyone can cook simultaneously, so the scheduler takes turns. Each cook's personal station with its board, knife, and spices is the stack: memory that belongs only to that thread. Walking to the storeroom to fetch a new cook is a system call: slow and costly. That's why serious restaurants don't hire and fire cooks mid-shift; they keep a fixed team and spread orders across them. That's exactly what an Executor does for you.
- Thread: an independent unit of execution. A thread drives one sequence of instructions.
- OS scheduler: the part of the operating system that decides which thread runs on which core and for how long. You don't control it directly; you only make requests.
- Stack: each thread's private memory for local variables and the chain of method calls. Every thread has its own (typically ~1 MB reserved).
- System call: when your program asks the OS kernel for something (e.g., "make a new thread"). These are relatively slow.
- Blocking: when a thread stops and waits for something to happen — e.g., waiting for a database reply. While blocked, that thread does nothing but wait.
- Monitor lock: a lock acquired on an object via the
synchronizedkeyword so only one thread at a time enters a critical section. Like a bathroom key only one person holds. - CPU-bound vs I/O-bound: CPU-bound work means the thread is truly computing (compression, hashing). I/O-bound work means the thread mostly waits on something external (DB, network, disk). This distinction is the backbone of pool sizing.
Now that we have the tools, let's meet the thread itself.
Mental model: what a thread is and why it's expensive
A java.lang.Thread in Java is really a thin wrapper over an OS thread — we call that OS-level thread a platform thread, to distinguish it later from the virtual thread.
Why do I say it's expensive? When you create a thread, three costly things happen:
- A large native stack is allocated (typically ~1 MB reserved).
- The thread is registered with the OS scheduler.
- All of this requires a system call.
That very expense of creating a thread is the entire reason the Executor framework exists. Instead of hiring a fresh cook for every order (slow and costly), you keep a small fixed team of cooks and spread work across them. This is called decoupling task submission from thread creation. You hand off the work; which ready thread runs it is no longer your concern.
Three questions frame this whole chapter. Answer these three and concurrency stops being mysterious:
- What is the unit of work?
Runnable(no result, no checked exception),Callable<V>(returns a value, may throw a checked exception), or the rawThreadsubclass (almost never the right answer). - Who runs it, and on what thread? A bare
Thread, or anExecutorServicebacked by a pool. - How does it stop? Cooperatively, via interruption — because Java has no safe preemptive
Thread.stop().
Let's unpack these three one by one.
Thread lifecycle & states
Over its life, a thread moves between several "states." Java keeps these in an enum called Thread.State. This enum is the ground truth — memorize it, because interviewers almost always probe the difference between WAITING and BLOCKED.
Picture an employee. NEW: just hired, hasn't started working yet. RUNNABLE: sitting at the desk, working or ready to work. BLOCKED: standing outside the conference room, waiting for the previous person to come out and hand over the key (the lock). WAITING: waiting for a colleague to signal them, with no deadline. TIMED_WAITING: same waiting, but with an alarm clock — "I'll wait until 3 o'clock." TERMINATED: done for the day and gone home.
| State | Meaning | How you get there |
|---|---|---|
NEW |
Created, not yet started | new Thread(...) before start() |
RUNNABLE |
Eligible to run or running (JVM does not distinguish "ready" vs "on-CPU") | after start() |
BLOCKED |
Waiting to acquire a monitor lock (synchronized) |
contending for a monitor |
WAITING |
Waiting indefinitely for another thread's signal | wait(), join(), LockSupport.park() |
TIMED_WAITING |
Same, but with a deadline | sleep(ms), wait(ms), join(ms), park(ns) |
TERMINATED |
run() returned or threw |
task finished |
Now two traps, exactly where the interviewer puts a finger:
This one surprises many people. A thread blocked on an I/O operation (say, stuck mid-read on a network socket, waiting for data) shows up in a thread dump as RUNNABLE, not BLOCKED. Why? Because that block happens at the OS level and the JVM cannot see it at all — from the JVM's point of view the thread is "able to run." Java uses the BLOCKED label only and specifically for contention over a synchronized monitor lock. So next time you see a thread dump full of RUNNABLE, don't be fooled; those threads may actually be waiting on the network.
start() is one-shot. If you call start() twice you get an IllegalThreadStateException. And crucially: if you call run() directly, no new thread starts at all — the body executes on the current thread, synchronously (right here, right now). Meaning no concurrency happens whatsoever. This is one of the most common accidental bugs beginners hit.
Runnable task = () -> System.out.println("on " + Thread.currentThread());
Thread t = new Thread(task);
t.run(); // BUG: runs on the *current* thread, no concurrency
t.start(); // correct: spawns a new thread
Keep the difference like this: run() is just an ordinary method that you call; it's start() that tells the OS "create a fresh thread and let that thread call run()."
Thread vs Runnable vs Callable/Future
Now the first question: what is the unit of work?
A golden rule: prefer composition over inheritance. That is, instead of extending Thread (extends Thread), write your work as a Runnable or Callable and hand it to a thread. Why? Because subclassing Thread couples your work to a thread forever and prevents you from reusing it in a pool. But a Runnable is a self-contained package of work that any thread can pick up and run.
Runnable/Callable is like an order ticket: on paper you've written "make one pizza." That ticket can be handed to any cook. But extends Thread is like building, for each order, a cook specific to that order with the work sewn into their body — useless for the next order. The order ticket is reusable; the sewn-in cook is not.
The difference between Runnable and Callable:
Runnablemeans "fire and forget": returns no value and its signature allows no checked exceptions.Callable<V>means "do something and give me the result back": returns a value of typeVand may throw a checked exception.
// Runnable: fire-and-forget, no return, no checked exceptions in signature
Runnable r = () -> log.info("done");
// Callable: returns a value, may throw checked exceptions
Callable<Integer> c = () -> {
return httpClient.fetchCount(); // may throw IOException
};
So if Callable returns a result but doesn't run right now — it finishes later on another thread — how do I get the result? That's where Future enters.
Future<V> is like the ticket the dry cleaner hands you. You've dropped off the work (the clothes), the result isn't ready yet, but you hold a handle you can use to collect the result later. When you call f.get(), it's like bringing the ticket back and saying "give me my clothes now" — and if they're not ready yet, you stand there and wait (that's what "get() is blocking" means).
ExecutorService pool = Executors.newFixedThreadPool(4);
Future<Integer> f = pool.submit(c);
try {
Integer n = f.get(2, TimeUnit.SECONDS); // blocks up to 2s
} catch (ExecutionException e) {
// the task threw; e.getCause() is the ORIGINAL exception
Throwable real = e.getCause();
} catch (TimeoutException e) {
f.cancel(true); // true => interrupt the running thread
}
Three things about Future that trip people up — read these carefully, because all three can create silent bugs in production:
This is Future's most dangerous behavior. If a task you handed off with submit() throws mid-work, nothing visibly happens. No log, no error, nothing. The exception is wrapped in an ExecutionException and delivered to you only when you call get() (with the original exception in e.getCause()). By contrast, if you hand off work with execute(Runnable) — which gives no Future at all — an uncaught exception goes to that thread's UncaughtExceptionHandler and usually gets printed. This asymmetry silently buries bugs. Rule: either always get() your Futures, or deliberately use execute.
f.cancel(true)only sets the interrupt flag; the task itself must cooperate (I'll explain in the interruption section below).cancel(false)prevents a not-yet-started task from running but won't interrupt a running one.get()is blocking. The calling thread stands there until the result is ready. If you want to chain work without standing still, preferCompletableFuture, which is non-blocking and composable (with methods likethenApply,thenCompose,allOf).
The Executor framework
Now the second question: who runs the work? The professional answer: an Executor.
The Executor framework has a clean hierarchy, going from simple to complete:
Executor— the most basic, has just one method:execute(Runnable).ExecutorService— extends it, addingsubmit,invokeAll, and lifecycle management (shutdown).ScheduledExecutorService— extends that further, adding delays and periodic execution.
But the real workhorse behind all of these is the class ThreadPoolExecutor. First, let's see why you shouldn't blindly trust the ready-made shortcuts.
The Executors factory — and why to distrust it
The Executors class has several convenient factory methods that build a pool for you in one line. The catch: several are dangerous in production, because they use either unbounded queues or unbounded thread counts:
Executors.newFixedThreadPool(n); // LinkedBlockingQueue with NO capacity bound → OOM risk
Executors.newSingleThreadExecutor();// same unbounded queue
Executors.newCachedThreadPool(); // maxPoolSize = Integer.MAX_VALUE → unbounded threads
Let me make the danger concrete:
"Unbounded" means "no ceiling." Under sustained heavy load, newFixedThreadPool grows its queue of waiting work without limit — work piles on work until the heap fills and you get an OutOfMemoryError. On the other side, newCachedThreadPool has no cap on thread count (Integer.MAX_VALUE), so under pressure it spawns threads until the OS refuses to make more and the app falls over. The senior practice is one sentence: construct ThreadPoolExecutor directly and by hand, with a bounded queue and an explicit rejection policy. That way, when the system is under pressure, instead of exploding it politely says "no."
ThreadPoolExecutor: the seven parameters
Let's look at the full constructor. It takes seven parameters, and each is an engineering decision:
ThreadPoolExecutor pool = new ThreadPoolExecutor(
4, // corePoolSize
16, // maximumPoolSize
60L, TimeUnit.SECONDS, // keepAliveTime for threads above core
new ArrayBlockingQueue<>(1000), // bounded work queue
new ThreadFactoryBuilder() // name threads! (Guava, or a hand-rolled factory)
.setNameFormat("api-worker-%d").build(),
new ThreadPoolExecutor.CallerRunsPolicy() // rejection handler
);
- corePoolSize (4): the number of always-kept threads, retained even when idle.
- maximumPoolSize (16): the most threads the pool is allowed to grow to.
- keepAliveTime (60s): threads above core that sit idle this long are torn down.
- work queue (ArrayBlockingQueue capacity 1000): where tasks wait for a free thread. Making this queue bounded is critical — you'll see why in a moment.
- ThreadFactory: the factory that creates threads; here we use it to name them.
- RejectedExecutionHandler: what to do with a new task when everything is full.
Imagine you run a café. corePoolSize is your permanent baristas (say 4). The queue is the waiting chairs (1000 of them). maximumPoolSize is the most baristas you can call up from the back (16). Now when a customer arrives: first, if you have fewer than your permanent baristas busy, one steps up. Then, if all are busy, the customer sits in a waiting chair. Only when all the chairs are full too do you call up an extra barista from the back. And if baristas have hit the max and the chairs are full, you tell the customer "no" at the door.
The submission algorithm of ThreadPoolExecutor is exactly this order and is worth memorizing, because it explains a notorious surprise:
- If running threads <
corePoolSize, start a new thread for the task (even if other threads are idle). - Else, try to enqueue the task.
- If the queue is full, and running threads <
maximumPoolSize, start a new (non-core) thread. - If the queue is full and threads == max, reject via the handler.
Here's the trap. If your queue is unbounded (like newFixedThreadPool), step 2 always succeeds — the queue never fills. So the "queue is full" condition in step 3 never becomes true, which means step 3 never fires and maximumPoolSize is completely dead. The pool never grows past corePoolSize, no matter how heavy the load. This surprises almost everyone. So if you want your pool to scale up under load, you must give it a bounded queue (or a SynchronousQueue, which has zero capacity — so every task either immediately demands a new thread or is rejected; that's exactly how newCachedThreadPool works).
Rejection policies (RejectedExecutionHandler)
When the pool is fully saturated (queue full and threads at max), what do we do with a new task? You have four ready-made policies:
| Policy | Behavior | Use when |
|---|---|---|
AbortPolicy (default) |
Throws RejectedExecutionException |
You want fail-fast and a caller that handles overflow |
CallerRunsPolicy |
Runs the task on the submitting thread | Backpressure: the submitter slows down because it's busy running the task, naturally throttling intake |
DiscardPolicy |
Silently drops the task | Only when losing work is acceptable (e.g., best-effort metrics) |
DiscardOldestPolicy |
Drops the oldest queued task, retries submit | Rarely correct; drops work that's been waiting longest |
"Backpressure" means that when the system is under strain, instead of blindly accepting more and more work until it bursts, it pushes back on the source: "send slower." CallerRunsPolicy gives you this for free: when the pool is full, the very thread that submitted the task is forced to run it. As long as it's busy running that task, it cannot submit new work — so the intake rate drops by itself. This is the pragmatic default for request pipelines. One warning though: if that submitting thread is the same thread accepting HTTP connections, running a task on it means accepting new connections stops for a while. Which may or may not be what you want.
Pool-sizing math
Now we reach where senior candidates truly earn their stripes: how many threads should I use? The answer depends on one thing — is your work CPU-bound or I/O-bound?
CPU-bound case: CPU-bound work (compression, hashing, pure computation) keeps a core busy the whole time. Here, if you make more threads than cores, you gain nothing; you only add context-switch overhead (the cost the OS pays to set one thread aside and put another on the core):
threads ≈ number_of_cores (or cores + 1 to cover the occasional page fault)
If you have 4 stoves and 4 cooks, one cook per stove, everything's perfect. Now bring 8 cooks to the same 4 stoves. Do you cook faster? No! There are still only 4 stoves. The extra cooks either stand idle or constantly swap places at the stove, and that swapping itself takes time. CPU-bound work means "the stove is always lit"; so peak speed is when the number of cooks equals the number of stoves.
I/O-bound case: I/O-bound work spends most of its time waiting (on the DB, HTTP, disk). Here the story is completely different: you want enough threads that while several are waiting, others can use the CPU. Brian Goetz's famous formula (from Java Concurrency in Practice), which is really a restatement of Little's Law:
threads = cores × targetUtilization × (1 + waitTime / computeTime)
Let's pin it with a worked example. Suppose 8 cores, target 100% utilization, and each task waits 90 ms on the DB and computes for 10 ms (so wait/compute = 9):
threads = 8 × 1.0 × (1 + 90/10) = 8 × 10 = 80 threads
See what happened? An I/O-heavy service legitimately wants about 80 threads on 8 cores — ten times the number of cores! The logic is simple: since each thread spends 90% of its time just waiting, you need many threads so that the cores never sit idle and one of them is always doing useful work. Undersize it and the CPU sits idle while requests queue; oversize it and you burn memory (~1 MB stack each) and thrash the system (constant, useless swapping). This wait/compute ratio is the single most important number in pool sizing — and it's exactly the pressure that virtual threads (at the end of the chapter) relieve.
int cores = Runtime.getRuntime().availableProcessors();
double waitToCompute = 9.0; // measured from profiling, not guessed
int size = (int) (cores * (1 + waitToCompute));
One important practical note: you must get the waitToCompute number from real profiling, not from a guess.
The formula gives you a number, but the real world has another hard ceiling. If your DB connection pool has only 20 connections, more than ~20 concurrent DB threads is useless — thread number 21 just queues behind the connection pool. So always bound your pool to the real downstream bottleneck. Measure where it actually gets narrow, and stop there.
ScheduledExecutorService
Sometimes you don't want the work to run right now; you want it "in 5 seconds" or "every 10 seconds." For these you have ScheduledExecutorService.
You may know the old Timer class. It's considered obsolete in practice, for two reasons: (1) it uses just one thread, so if one task is slow the rest fall behind; and (2) if a task throws an exception, the whole Timer silently dies. ScheduledExecutorService is its proper, resilient replacement.
ScheduledExecutorService sched = Executors.newScheduledThreadPool(2);
sched.schedule(() -> log.info("once, after 5s"), 5, TimeUnit.SECONDS);
// fixed RATE: fires every 10s regardless of task duration (can pile up / overlap-guarded)
sched.scheduleAtFixedRate(this::poll, 0, 10, TimeUnit.SECONDS);
// fixed DELAY: waits 10s AFTER each completion — no pile-up
sched.scheduleWithFixedDelay(this::poll, 0, 10, TimeUnit.SECONDS);
You should feel the difference between these two methods in your bones, because it's a critical distinction:
scheduleAtFixedRate is like a train whose departure time is fixed: it leaves on the hour every hour, no matter how long the previous trip took. If one run takes longer than the 10-second period, the next run starts immediately after it to catch up (but they never run concurrently on the same task). scheduleWithFixedDelay, by contrast, is like a bus that rests for 10 seconds after returning to the depot each time, then sets off again — so the gap is always between the end of one and the start of the next, and runs never pile up.
scheduleAtFixedRate: targets a fixed period. If a run exceeds the period, subsequent runs execute back-to-back to catch up.scheduleWithFixedDelay: inserts a fixed gap between the end of one run and the start of the next, so runs never overlap.
This one can genuinely eat hours of your life. If a scheduled task throws an uncaught exception, that task is cancelled and never runs again — the pool silently suppresses it and no further executions happen, with no error in the logs. Picture a cron that ran for months and one day "just stops." This is almost always the cause. Always wrap the body of periodic tasks in try/catch:
sched.scheduleWithFixedDelay(() -> {
try { doWork(); }
catch (Throwable t) { log.error("poll failed, will retry next tick", t); }
}, 0, 10, TimeUnit.SECONDS);
Shutdown: shutdown() vs shutdownNow()
An important note: you must shut down an ExecutorService yourself, or your JVM won't exit at all! Because pool worker threads are by default non-daemon (user) threads, and while alive they keep the JVM alive. (I'll explain daemon below.)
You have two ways to shut down:
| Method | Accepts new tasks? | Already-queued tasks? | Running tasks? |
|---|---|---|---|
shutdown() |
No | Run to completion | Finish |
shutdownNow() |
No | Returned as List<Runnable>, not run |
Interrupted |
shutdown() means "we lock the front door, take no new customers, but everyone already inside who ordered gets their meal in full." shutdownNow() means "closed right now! Cooks, drop what you're doing (interrupt), and orders not yet started stay on the counter" — it hands those unstarted orders back to you as a list so you can decide what to do with them.
Note that shutdownNow() interrupts running tasks (sets their interrupt flag) but cannot force-stop a task that ignores interruption. The canonical, correct graceful-shutdown idiom is this:
void shutdownGracefully(ExecutorService pool) {
pool.shutdown(); // stop accepting, let queued tasks finish
try {
if (!pool.awaitTermination(30, TimeUnit.SECONDS)) {
pool.shutdownNow(); // force: interrupt running tasks
if (!pool.awaitTermination(10, TimeUnit.SECONDS))
log.error("pool did not terminate");
}
} catch (InterruptedException e) {
pool.shutdownNow();
Thread.currentThread().interrupt(); // restore the flag
}
}
Two subtle points: first, awaitTermination returns only a boolean and does not itself trigger shutdown — you must call shutdown() or shutdownNow() first, otherwise you'll wait forever. Second, since Java 19, ExecutorService implements AutoCloseable; its close() method behaves like "shutdown() then await," so you can put the pool in a try-with-resources and rest easy that it closes automatically.
The interruption model (cooperative cancellation)
We've reached the third question: how does a thread stop? And here Java made a deliberate, important decision: there is no safe way to forcibly kill a thread.
You might ask, "don't we have a Thread.stop() method?" We do, but it's deprecated and you must never use it. The reason is dangerous: stop() cuts the thread off mid-work and releases its locks mid-mutation — that is, in the middle of a half-finished change to an object. The result? Objects left in a corrupt, inconsistent state. Like yanking a surgeon out of the room mid-operation. That's why Java closed this door.
So cancellation in Java is cooperative. You only request cancellation (via interrupt()), and the target thread itself chooses to notice and cleanly wind down.
interrupt() is like gently tapping a colleague who's wearing headphones on the shoulder and saying "whenever you can, finish up and come over." You can't switch off their brain; you only send a signal. If they're cooperative, they notice at the first opportunity and wind down cleanly. If they're oblivious and never check, nothing happens. Cooperative cancellation is exactly this: it needs the target thread's cooperation.
Interruption is really just a per-thread boolean flag, plus a few rules:
thread.interrupt()— sets the flag (sends the signal).Thread.interrupted()— static; reads and clears the flag on the current thread.thread.isInterrupted()— reads the flag without clearing it.- Blocking methods that declare
throws InterruptedExceptionin their signature (sleep,wait,join,BlockingQueue.take,Future.get) clear the flag and throw anInterruptedExceptionwhen interrupted.
Now, when you catch an InterruptedException, you have exactly two correct responses:
- Propagate it (let it bubble up to higher layers), or
- Restore the flag and exit the work, so callers up the stack can still see the cancellation:
try {
queue.take();
} catch (InterruptedException e) {
Thread.currentThread().interrupt(); // NEVER swallow silently
return; // stop the unit of work
}
The worst thing you can do is catch InterruptedException with an empty catch and do nothing. Remember: when this exception is thrown, the interrupt flag has already been cleared. Now if you swallow it too, you've completely destroyed the one signal that told the world "stop" — and now your task can no longer be cancelled, and even shutdownNow can't stop it. This is one of the most-flagged bugs in senior code review. Never, ever swallow it silently.
One last note: blocking methods throw InterruptedException themselves, but if you have a CPU-bound loop that calls no blocking method, you must check the flag manually:
while (!Thread.currentThread().isInterrupted()) {
doOneChunk();
}
// loop exits cleanly on interrupt — cooperative cancellation
ThreadLocal — and why it leaks in pools
ThreadLocal<T> gives each thread its own private copy of a variable. It's the standard mechanism for holding "per-request context" — like a transaction ID, the current user, or a SimpleDateFormat instance — without having to pass it by hand through every method.
ThreadLocal is like each employee having a personal locker at work. Everyone has access to "the locker" (the variable is shared), but the inside of each person's locker is theirs alone and no one sees another's contents. This is great, until… the employees don't change. If you later give that same locker to a new employee without emptying it, the previous person's belongings are still inside. And that is exactly what happens in a thread pool.
private static final ThreadLocal<SimpleDateFormat> FMT =
ThreadLocal.withInitial(() -> new SimpleDateFormat("yyyy-MM-dd"));
ThreadLocal values are stored in a map inside the Thread object itself, keyed by the ThreadLocal object. In pooled threads this creates two serious hazards:
Pool threads are long-lived and reused. A value you set while serving request A is still there when the same thread later serves request B. Now request B sees request A's context — for example, user A's identity! This is both a correctness bug and a serious security bug. The iron rule: if you set a ThreadLocal in a pooled thread, you must remove() it in a finally block:
try {
RequestContext.CURRENT.set(ctx);
handle(request);
} finally {
RequestContext.CURRENT.remove(); // mandatory in pools
}
ThreadLocal's internal map (the ThreadLocalMap) uses weak references for its keys (the ThreadLocal object itself) but strong references for the values. Now think: if the ThreadLocal object is no longer referenced anywhere but the thread stays alive (which is exactly how pool threads behave), the value can't be collected by the garbage collector — until that map slot is reused or the thread dies. If you've cached large objects in a ThreadLocal and your pool is long-lived, the heap slowly fills. remove() fixes this too.
You might think InheritableThreadLocal is the answer — this variant copies the value to a child thread at creation time. But pool threads aren't re-created per task (that's the whole point of a pool!), so it does not propagate context into pool tasks. To propagate context correctly in pools you need explicit wrapping, or ScopedValue, which arrived in Java 21.
Daemon threads
Every thread in Java is either a user thread or a daemon thread. The key rule: the JVM exits precisely when the last non-daemon thread finishes — and at that moment, it abruptly abandons any daemon thread still running (with no guarantee that finally blocks run).
A user thread is like a permanent employee: as long as one has work, the building (the JVM) stays open. A daemon thread is like a night watchman or the HVAC system: it does background work, but the moment the last permanent employee leaves, the lights go out and the watchman too — like it or not, even mid-task — is left outside. That's why you never assign critical work to the night watchman.
Thread t = new Thread(task);
t.setDaemon(true); // must be set BEFORE start(), else IllegalThreadStateException
t.start();
- Use daemon threads for background infrastructure that should never keep the app alive: metrics flushers, cache-eviction timers, heartbeat pollers.
- Never use a daemon thread for work that must complete or must clean up (writing a file, flushing a transaction) — the JVM may kill it mid-write.
- A thread inherits its parent's daemon status by default. That's why
Executorsthreads are user (non-daemon) threads and keep the JVM alive until you shut the pool down. If you want daemon threads in a pool, you must supply aThreadFactorythat callssetDaemon(true).
Virtual threads (Java 21+) — where the sizing math evaporates
And now the biggest change of recent years. Java 21 made virtual threads permanent. Let me build the intuition first.
Imagine you run a courier company. A platform thread is like an expensive company car: there are only a few, and each is costly. If a courier reaches a door and has to wait for it to open, that expensive car sits parked and idle right there. A virtual thread is like a cheap courier: when they reach a door and must wait, they get off the car (unmount) and free it for another courier to hop on; when the door opens, they mount a free car again. There are few cars (the platform carrier threads = carriers), but there can be millions of couriers.
So a virtual thread is a thread that the JVM itself schedules onto a small pool of carrier platform threads; when it blocks on I/O, instead of blocking an OS thread it unmounts — it parks in user space and frees its carrier. They're cheap enough that you can create millions.
// Each task gets its own virtual thread; NOT a pool to be sized
try (var exec = Executors.newVirtualThreadPerTaskExecutor()) {
for (var req : requests)
exec.submit(() -> handleBlocking(req)); // blocking calls are fine here
} // close() awaits all tasks (AutoCloseable)
This is the heart of it. For I/O-bound work, you no longer size a pool. You write straightforward blocking code, one virtual thread per task, and let the JVM multiplex them onto the few carriers. Remember Goetz's formula? That formula was really a workaround for the "expensive platform thread" problem. Virtual threads remove, at the root, the very constraint the formula was compensating for.
But a few things still bite:
If a virtual thread is inside a synchronized block (or in the middle of a native/JNI call) and blocks there, it can no longer unmount from its carrier — it gets pinned to that carrier, i.e., nailed to it. If enough threads get pinned to the carriers, the carrier pool starves and throughput collapses. The fix: around blocking I/O, use a ReentrantLock instead of synchronized, which lets the virtual thread unmount. (Later JDKs, 24+, remove most synchronized-induced pinning, but on Java 21 itself the problem is real.)
- Don't pool them. Making a
newFixedThreadPoolof virtual threads defeats the whole philosophy; always create one per task. - They give no benefit for CPU-bound work. A virtual thread helps with waiting, not computing. For pure computation, keep the platform-thread pool sized to your cores.
Common pitfalls & gotchas (rapid fire)
Let's review all the traps we saw in this chapter in one place — these are exactly the ones that cause pain in production:
- Calling
run()instead ofstart()— no new thread. - Unbounded
newFixedThreadPoolqueue → OOM under load. maximumPoolSizeignored because the queue is unbounded.- Swallowing
InterruptedException→ uncancellable tasks. - Forgetting
ThreadLocal.remove()in a pool → leaked/stale context. - Not naming threads → useless thread dumps in production. Always supply a
ThreadFactory. - Exceptions in
submitted tasks silently lost untilget(). - Scheduled task throws once → silently stops forever.
- Sharing a non-thread-safe object (
SimpleDateFormat,Randomunder contention) across pool threads.
Best practices
- Build
ThreadPoolExecutorexplicitly with a bounded queue + a chosen rejection policy; avoid theExecutorsunbounded factories in production. - Name your threads via
ThreadFactory— future-you reading a thread dump will thank you. - Size CPU pools to
cores; size I/O pools with the wait/compute formula, then cap at the real downstream bottleneck. - Always
shutdown()+awaitTermination()+shutdownNow()on the way down. - Treat
InterruptedExceptionas a first-class signal: propagate or restore-and-exit, never swallow. remove()everyThreadLocalin afinallywhen threads are pooled.- Prefer
CompletableFuture/structured concurrency over chains of blockingget(). - On Java 21+, use virtual threads for I/O-bound fan-out and stop hand-tuning pool sizes.
Interview Questions
Now let's practice everything you've learned in the shape of real interview questions. Answer each yourself first, then read the answer.
BLOCKED means the thread is contending specifically for a synchronized monitor lock. WAITING/TIMED_WAITING mean it's parked awaiting a signal. RUNNABLE means the JVM considers it able to run. The JVM cannot see OS-level I/O blocking, so a thread stuck in a blocking socket read is still labeled RUNNABLE — the block happens below the JVM's visibility.
A second start() throws IllegalThreadStateException — a thread is one-shot. Calling run() executes the body synchronously on the current thread; no new thread is created. This is a common accidental bug.
Only 4. With an unbounded queue, tasks are always enqueued successfully, so the pool never reaches the "queue full" condition that triggers creating threads above core. maximumPoolSize is effectively dead. To use max, you need a bounded queue.
(1) If threads < core, create a new thread. (2) Else try to enqueue. (3) If enqueue fails (queue full) and threads < max, create a non-core thread. (4) Else invoke the rejection handler. The subtle point: enqueue is attempted before growing past core, which is why an unbounded queue caps you at core.
AbortPolicy throws (fail-fast, default). DiscardPolicy drops silently. DiscardOldestPolicy drops the head of the queue and retries. CallerRunsPolicy runs the task on the submitting thread — this provides backpressure because the submitter is busy and can't submit more, throttling intake.
threads = cores × util × (1 + wait/compute) = 16 × 1.0 × (1 + 190/10) = 16 × 20 = 320. But also cap at the downstream's real concurrency limit — 320 threads hammering a service that allows 50 concurrent just queues; size to the bottleneck.
The interrupt flag is cleared when the exception is thrown. If you catch it and do nothing, you've destroyed the only cancellation signal — the task becomes uninterruptible and shutdownNow can't stop it. Correct: either rethrow/propagate it, or call Thread.currentThread().interrupt() to restore the flag and then exit the work.
shutdown(): stops accepting new tasks, lets queued and running tasks finish. shutdownNow(): stops accepting, returns un-started queued tasks as a List<Runnable> (they never run), and interrupts running tasks (which only stops them if they honor interruption).
Correctness: pool threads are reused, so a value set for one task persists into the next task on the same thread unless you remove() it — leaking one request's context (e.g., user identity) into another. Memory: ThreadLocalMap keys are weak but values are strong; a long-lived pool thread holds the value even after the ThreadLocal is gone, until the slot is reused. remove() in a finally fixes both.
ExecutorService pool = Executors.newSingleThreadExecutor();
Future<?> f = pool.submit(() -> { throw new RuntimeException("boom"); });
System.out.println("submitted");
// no f.get() call
pool.shutdown();
It prints submitted and nothing about "boom". Exceptions from submitted tasks are captured in the Future and surface only on f.get() (as ExecutionException). Without get(), the exception is silently swallowed. (Had you used execute(), it would reach the UncaughtExceptionHandler and print a stack trace.)
scheduleAtFixedRate targets a fixed period; if a run exceeds the period, the next run starts immediately after (bursting to catch up, never overlapping the same task). scheduleWithFixedDelay always inserts the full delay after completion, so runs never pile up. For polling that must not overlap or stampede, use fixed delay.
It threw an uncaught exception. A scheduled task that throws is silently cancelled and never scheduled again; the exception is trapped in the (ignored) Future. Wrap the task body in try/catch to keep it alive.
Daemon status is immutable once the thread is alive — setting it after start() throws IllegalThreadStateException. A daemon loses data when it's the only thing running a critical operation and all user threads exit: the JVM terminates and abandons the daemon mid-operation, without guaranteeing finally/cleanup runs. Never put must-complete work on daemons.
Pinning. The blocking I/O sits inside a synchronized block (or a native call), so the virtual thread can't unmount from its carrier while blocked. With few carriers, pinned virtual threads starve the pool. Fix: replace synchronized around the blocking call with a ReentrantLock, which lets the virtual thread unmount. (Also verify you're using per-task virtual threads, not pooling them.)
For CPU-bound work — they don't add parallelism beyond the carrier count, which is bounded by cores. Compute-heavy tasks should use a platform-thread pool sized to cores. Virtual threads only win when threads spend time waiting.
Senior notes & advanced edge cases
You now hold the mental model of threads, Executors, pool sizing, cooperative cancellation, and virtual threads. Now we go one layer deeper — the layer you only learn after a few on-call nights and a couple of real post-mortems: the things that look correct in code and detonate under production load. This section repeats none of the earlier material; it only fills the gaps.
- The Java Memory Model: the happens-before you get for free from an Executor, and why you don't need
volatileto read a task's result. availableProcessors()lies inside containers — the biggest sizing mistake in Kubernetes.ForkJoinPool.commonPool()and parallel streams share one pool, and blocking in it starves the whole JVM.CompletableFuturedone right — which thread runs the callback,joinvsget, collecting results.- Thread-starvation deadlock (a pool nested in itself).
- Observability and hooks, context/MDC propagation, queue types.
- Spring
@Asynctraps. - The real virtual-threads playbook: a Semaphore instead of sizing, pinning in modern JDKs, ScopedValue, structured concurrency.
- Seven hard senior interview questions with full answers.
The memory model: happens-before you get for free
A subtle question: a task on thread B computes a value and writes it into an ordinary field (no volatile); thread A reads that field after future.get(). Is A guaranteed to see the fresh value, or might it see a stale/half-built one?
The java.util.concurrent package spec explicitly guarantees two happens-before edges:
- Everything you did on the submitting thread before
submit/executehappens-before the task begins. The task sees all your pre-submit writes. - Everything the task does happens-before
Future.get()returns. Afterget(), thread A is guaranteed to see all of the task's writes, fully. So you do not needvolatileorsynchronizedto hand a task's result across threads — the submit→run→get boundary establishes the visibility itself. The same guarantee holds forBlockingQueue(put/take),CountDownLatch, and a thread's termination relative tojoin().
If you execute() a task and then read its result from another thread with no synchronization mechanism at all (no get, no queue, no latch, no volatile field), you have no happens-before edge and may see a stale value — or a half-constructed object — forever. The bug is non-deterministic and invisible in tests.
availableProcessors() lies inside containers
Every sizing formula from the previous chapter leans on "the number of cores." But where does that number come from? Runtime.getRuntime().availableProcessors(). The problem: in Kubernetes that number is often not what you think.
Your node has 32 cores, but you gave the pod a cpu limit of, say, 500m (half a core). Since JDK 10 the JVM is container-aware and reads the cgroup; with a sub-core quota, availableProcessors() returns 1. Now three disasters:
- If you sized a pool at a hard-coded 32, those 32 threads context-switch endlessly over a one-core quota and you get throttled.
ForkJoinPool.commonPool(), whose parallelism isavailableProcessors()-1, effectively becomes 0 (one thread), so every parallel stream silently runs serially.- Conversely, without a limit but with an old/misconfigured JVM, it may see 32 and trample every neighbor on the node.
Log the real number at startup (log.info("cores={}", Runtime.getRuntime().availableProcessors())) — that one line erases hours of debugging. If a fractional limit forces a wrong count, set -XX:ActiveProcessorCount=N explicitly, and prefer cpu request == cpu limit for predictable behavior. (cgroups v2 is fully supported only from JDK 17 onward.)
commonPool() and parallel streams share one pool
stream().parallel() and CompletableFuture.supplyAsync(task) (the no-executor overload) both run on one global pool: ForkJoinPool.commonPool(). That pool is small (core-count) and built for short CPU-bound work.
If you put a blocking call (HTTP, DB, Thread.sleep) inside a parallelStream().map(...), that worker holds a commonPool thread hostage. Because this pool is shared across the entire JVM, every other part of the app that relies on a parallel stream or supplyAsync now starves, and unrelated requests slow down. The classic incident: one reporting endpoint using a parallel stream drags the whole service to its knees.
Two correct fixes:
// Fix 1: run the parallel stream inside a dedicated ForkJoinPool
ForkJoinPool custom = new ForkJoinPool(16);
List<R> out = custom.submit(() ->
items.parallelStream().map(this::blockingCall).toList()
).get();
// Fix 2 (for CompletableFuture): always pass your own executor
CompletableFuture.supplyAsync(this::blockingCall, myIoPool);
If you must block inside an FJP, use ForkJoinPool.managedBlock(...) instead of a plain block. This interface tells the pool "I'm about to block," and the pool temporarily spins up a compensating worker to preserve parallelism. CompletableFuture itself uses this ManagedBlocker mechanism for join() when called on an FJP worker thread.
CompletableFuture done right: where the callback runs
The chapter introduced CompletableFuture as "non-blocking and composable." The senior point is one question: which thread runs the callback?
- The non-
Asyncform (e.g.thenApply): the callback runs on whichever thread completed the previous stage — or even on the calling thread if the stage was already complete. Unpredictable and dangerous for heavy work. - The
Asyncform with no executor (thenApplyAsync): oncommonPool()— with the starvation trap above. - The
Asyncform with an executor: on your own pool. This is the correct production default.
CompletableFuture
.supplyAsync(() -> fetchUser(id), ioPool) // I/O on our pool
.thenApplyAsync(this::enrich, cpuPool) // CPU on a separate pool
.exceptionally(ex -> User.anonymous()) // fallback
.thenAccept(this::send);
If your function itself returns a CompletableFuture, thenApply leaves you with a nested CompletableFuture<CompletableFuture<T>>; use thenCompose (the flatMap of futures) to flatten. get() throws checked exceptions; join() is the unchecked version (CompletionException) and is cleaner inside chains. To gather many futures use allOf(a,b,c).thenApply(...) — but allOf returns Void, so join each one afterward. For errors: exceptionally catches only the error path, handle catches both, and whenComplete is a finally that observes but does not transform.
Thread-starvation deadlock: a pool nested in itself
This is one of the sneakiest deadlocks and never shows up in tests, because it only occurs under load.
The sequence: your pool is bounded (say 10 threads). A task running on that pool submits subtasks to the same pool and then get()s, waiting on their results. If 10 parent tasks simultaneously occupy all 10 threads and each waits on its own subtask, no thread is left to run any subtask — everyone waits on everyone; total lock.
Diagram — why a bounded, self-nested pool deadlocks:
sequenceDiagram
participant P as Parent tasks (fill all N threads)
participant Q as Work queue
participant C as Child tasks
P->>Q: submit child, then block on get()
Q-->>C: waiting for a free thread
Note over P,C: all N threads held by parents<br/>no thread free to run children
Note over P,C: parents wait children, children wait threads = deadlock
A task running on a pool must never submit work to the same pool and then block on its result. Either use separate pools for separate layers, write the whole chain non-blocking with thenCompose, or use virtual threads (where threads are not scarce, which eliminates this whole class of deadlock).
Observability and hooks
A pool you can't see will ambush you one day. Feed getActiveCount(), getQueue().size(), and getCompletedTaskCount() to Micrometer — a perpetually full queue means the pool is too small or the downstream is slow. And override ThreadPoolExecutor's beforeExecute/afterExecute: the perfect place to set/clear MDC, time execution, and catch the exceptions that submit silently swallows (afterExecute receives both the Runnable and the Throwable). Two more knobs: allowCoreThreadTimeOut(true) lets idle core threads be reclaimed (low-traffic pools), and prestartAllCoreThreads() warms the lazily-created core threads up front when first-request latency matters.
SLF4J MDC ThreadLocals, SecurityContextHolder, and tracing trace-ids are not automatically carried onto pool threads — the worker doesn't have them. Fix: a TaskDecorator (in Spring) or a wrapper that copies the context from the submitting thread before running and clears it after. Virtual threads ease this pain but don't remove it.
Choosing the queue type is an engineering decision too: ArrayBlockingQueue is the safe production default, PriorityBlockingQueue for a prioritized queue, and DelayQueue for scheduling — all in preference to the unbounded, dangerous LinkedBlockingQueue.
Spring @Async traps
@Async is the prettiest "go async" in the world — until one of these bites you:
- Self-invocation:
@Asyncworks via a proxy. If a public method calls the@Asyncmethod on the same bean, the call never crosses the proxy and runs completely synchronously — with no error. You must invoke it from a different bean. - Lost exceptions: if the
@Asyncmethod returnsvoidand throws, the exception goes nowhere (just likeexecutewithout a Future) and is lost from the logs. To catch it, define anAsyncUncaughtExceptionHandler, or make the return type aCompletableFutureso the error surfaces onget. - The default executor: in plain Spring (no Boot), if you don't specify an executor the fallback is
SimpleAsyncTaskExecutor, which has no pool and spins up a fresh thread for every call — thread explosion under load. Spring Boot mercifully auto-configures aThreadPoolTaskExecutor, but don't trust even that default — configure it explicitly.
Since Spring Boot 3.2, setting spring.threads.virtual.enabled=true (on Java 21+) switches both the @Async executor and Tomcat's request threads to virtual threads, so your traditional blocking code scales without a rewrite. But the same MDC and pinning caveats still apply.
The real virtual-threads playbook
The chapter said that with virtual threads you no longer size a pool. The senior's next question: so how do I protect a constrained downstream?
When you spawn one virtual thread per task, the pool is no longer a boundary — but your DB may allow only 20 connections. If 10,000 virtual threads hit it at once, the connection pool blows up or times out. The fix: bound concurrency not with pool size, but with a Semaphore around the scarce resource:
Semaphore db = new Semaphore(20);
db.acquire();
try { return jdbc.query(...); }
finally { db.release(); }
In the virtual-threads world, the Semaphore plays exactly the role maximumPoolSize used to play: a guard on downstream capacity.
With platform threads a common trick was to cache an expensive object (a big buffer, say) in a ThreadLocal, because threads were few. With millions of virtual threads this is an anti-pattern: one copy per thread means explosive memory use. With virtual threads, either build the object per-task, pool it, or consider ScopedValue.
The chapter said synchronized causes pinning. JEP 491 in JDK 24 fixed exactly that: a synchronized block no longer normally holds the carrier hostage. Only native/JNI and a few rare corners remain. To monitor pinning, use the JFR event jdk.VirtualThreadPinned instead of the old -Djdk.tracePinnedThreads flag (deprecated in JDK 24). You tune carrier count with -Djdk.virtualThreadScheduler.parallelism.
Build one manually with Thread.ofVirtual().name("job-", 0).start(runnable) or Thread.startVirtualThread(r). Two modern companions round out the fan-out pattern: structured concurrency (still a preview in JDK 25, JEP 505, now via StructuredTaskScope.open() with a Joiner) launches several subtasks in one scope so they all fail/cancel together and never leak a thread; and ScopedValue (finalized in JDK 25, JEP 506) is the immutable, cheaper ThreadLocal replacement for passing context to subtasks — no remove() needed, since its scope closes automatically.
Hard senior interview questions
No. The JMM guarantees that everything the task does happens-before Future.get() returns, and everything you wrote before submit happens-before the task starts. So the submit→run→get boundary establishes full visibility; no volatile/synchronized is needed to hand the result across. But if you read the result from a fire-and-forget task with no synchronization point (no get, no queue, no latch), you have no guarantee and a stale value is possible.
Because since JDK 10 the JVM is container-aware and reads the cgroup: with a one-core limit, availableProcessors() returns 1, not 32. So your pool is 2 threads, and worse, commonPool() (parallelism cores-1) becomes single-threaded, so parallel streams run serially. Fix: log the real number, set -XX:ActiveProcessorCount explicitly if needed, and prefer cpu request == limit.
Because a parallel stream runs on the global ForkJoinPool.commonPool(), and the blocking remote call holds its threads hostage; since that pool is shared JVM-wide, every other user of it starves. Fix 1: submit the stream inside a dedicated ForkJoinPool. Fix 2 (better): don't use a parallel stream for blocking work — use CompletableFuture.supplyAsync(task, dedicatedIoPool) or virtual threads. If you must block inside an FJP, use a ManagedBlocker so the pool spins up a compensating worker.
Thread-starvation deadlock. If 10 parent tasks simultaneously hold all 10 threads and each waits on its own subtask, no thread is left to run any subtask — parents wait on children, children wait on a free thread. Fix: separate pools per layer, a non-blocking chain with thenCompose, or virtual threads, where threads aren't scarce.
thenApply: the callback runs on whichever thread completed the previous stage (or even the calling thread if already complete) — unpredictable for heavy work. thenApplyAsync with no executor: on commonPool (the starvation trap). thenApplyAsync with an executor: on your pool — the correct production form. get() throws checked exceptions (ExecutionException); join() is unchecked (CompletionException) and cleaner inside chains.
Because @Async works via a proxy; a self-invocation never crosses the proxy, so no async wrapping happens and the code runs synchronously — you must call it from another bean or self-inject the proxy. On exceptions: an @Async method returning void is like execute without a Future; the exception is caught nowhere and lost from the logs. Fix: define an AsyncUncaughtExceptionHandler or return a CompletableFuture so the error surfaces on join/get.
You bound concurrency with a Semaphore(20) around the downstream call; in the virtual-threads world the Semaphore plays the role maximumPoolSize used to. For pinning: in JDK 24, JEP 491 means synchronized usually no longer pins; only native/JNI remains. To monitor, use the JFR event jdk.VirtualThreadPinned (the old -Djdk.tracePinnedThreads flag was deprecated in JDK 24).
submit → run → Future.getestablishes happens-before; no volatile needed to hand a result across, but fire-and-forget with no sync point is unsafe.- Inside a container,
availableProcessors()comes from the cgroup and may be far below the node's cores — breaking every sizing formula. commonPool()and parallel streams are global; blocking in them starves the whole JVM. Give a dedicated executor or use a ManagedBlocker. And a bounded self-nested pool is a thread-starvation deadlock.- In
CompletableFutureuse theAsyncform with your own executor; make the pool observable and propagate MDC with a decorator. @Asynctraps: self-invocation goes synchronous,voidswallows exceptions, the default may have no pool.- With virtual threads: a Semaphore replaces sizing, no expensive ThreadLocals, pinning is nearly solved in JDK 24; ScopedValue and structured concurrency are the modern fan-out pattern.
- A platform thread is expensive (~1 MB stack + a system call); that's why Executors decouple task submission from thread creation.
- Memorize the six
Thread.Statevalues.BLOCKEDis only monitor contention; OS-level I/O shows asRUNNABLE. run()creates no thread,start()does.Runnablereturns nothing,Callabledoes and you fetch it withFuture.get()— which is blocking and hides exceptions until the moment ofget().- In production, build
ThreadPoolExecutorby hand with a bounded queue and a rejection policy. With an unbounded queue,maximumPoolSizeis inert. - Size CPU-bound to core count; size I/O-bound with
cores × (1 + wait/compute), and cap at the real bottleneck. - For scheduling, know fixed-rate vs fixed-delay and wrap the task body in try/catch (else silent death).
- Always
shutdown()→awaitTermination()→shutdownNow(). Cancellation is cooperative; never swallowInterruptedException. - In pools,
remove()everyThreadLocalin afinally(correctness + security + memory). - On Java 21+, for I/O-bound work create one virtual thread per task and forget sizing — just watch out for pinning and don't pool them.
Sources: ThreadPoolExecutor (JDK 17), Managing Throughput with Virtual Threads — inside.java, RejectedExecutionHandler — Baeldung