Java Core · جاوا پایه سنیورSenior ~65 دقیقه مطالعه~56 min read
درونیات JVM، حافظه و Garbage CollectionJVM Internals, Memory & Garbage Collection
از صفر: JVM چیست، حافظه کجا زندگی میکند، GC چطور تمیز میکند و JIT چطور سریع میکند — با تشبیههای واقعی، کدِ اجراشدنی و سؤالات مصاحبه.From zero: what the JVM is, where memory lives, how GC cleans up and JIT speeds things up — with real-world analogies, runnable code, and interview questions.
این فصل، قلب جاواست. اگر این را خوب بفهمی، بقیهٔ زبان برایت «منطقی» میشود نه «حفظی». پس عجله نکن؛ قدمبهقدم میرویم — اول با یک تشبیه ساده مفهوم را میگیریم، بعد واژهٔ فنیاش را میبینیم، و آخر سرش وصل میکنیم به کد واقعی و سؤال مصاحبه.
اول میفهمیم JVM چیست و چرا اصلاً وجود دارد. بعد سه بخش بزرگش را باز میکنیم: (۱) چطور کلاسها بارگذاری میشوند، (۲) حافظه کجا زندگی میکند (heap، stack، metaspace)، (۳) موتور اجرا چطور کد را سریع میکند (JIT) و حافظه را تمیز میکند (Garbage Collector). در آخر میرسیم به تیونینگ، نشتی حافظه، خواندن لاگ GC و سؤالات مصاحبه.
بخش ۰ — سه واژهای که باید از صفر بفهمی
قبل از هر چیز، سه کلمهای که در کل فصل تکرار میشوند را با زبان آدمیزاد باز کنیم. اگر اینها جا بیفتند، بقیه راحت است.
«انتزاع» (Abstraction) یعنی چه؟
فرض کن رانندگی میکنی. تو فقط با فرمان، پدال گاز و ترمز کار داری. اینکه داخل موتور چه انفجارهایی میافتد، بنزین چطور میسوزد و چرخدندهها چطور میچرخند، اصلاً به تو ربطی ندارد و لازم نیست بدانی.
این یعنی انتزاع: «پیچیدگی را قایم کن، فقط چیزی را نشان بده که کاربر لازم دارد.»
«ماشین مجازی» (Virtual Machine) یعنی چه؟
JVM یک کامپیوترِ فرضی و ساختگی است که داخل کامپیوتر واقعی تو (ویندوز/مک/لینوکس) اجرا میشود. تو کدت را برای این «کامپیوتر خیالی» مینویسی، نه برای ویندوز یا مک. JVM پیچیدگیهای سختافزار واقعی را انتزاع (قایم) میکند و کد تو را به زبان سختافزار واقعی ترجمه میکند.
نتیجه: یک بار کد بنویس، همهجا اجرا کن — روی ویندوز، مک، لینوکس، سرور، موبایل. این همان شعار معروف جاوا است: Write once, run anywhere.
پس وقتی میگوییم JVM یک «ماشینِ انتزاعی» است، یعنی یک کامپیوتر خیالی که جزئیات سختافزار زیرش را از تو پنهان کرده.
«پشتهمحور» (Stack-based) یعنی چه؟
JVM برای انجام محاسبات از یک ساختار به اسم پشته (stack) استفاده میکند. پشته یعنی دقیقاً همان چیزی که از اسمش پیداست: یک دستهٔ بشقاب. آخرین بشقابی که میگذاری، اولین بشقابی است که برمیداری (به این میگویند LIFO: آخرینورودی، اولینخروجی).
مثلاً برای محاسبهٔ 2 + 3، JVM اینطور عمل میکند: 2 را روی پشته میگذارد، 3 را رویاش میگذارد، بعد دستور add را میبیند، هر دو را برمیدارد، جمع میکند و 5 را روی پشته میگذارد. لازم نیست الان جزئیاتش را حفظ کنی؛ فقط بدان پشته میز کارِ لحظهایِ JVM است.
انتزاع = قایمکردن پیچیدگی. ماشین مجازی = کامپیوتر خیالی که کد جاوا را اجرا میکند تا به سختافزار وابسته نباشی. پشتهمحور = میز کارِ لحظهای برای محاسبات، با منطق «آخرینورودی، اولینخروجی».
بخش ۱ — JVM دقیقاً چه میکند؟ (تصویر بزرگ)
بیا کل مسیر از «کدی که تو مینویسی» تا «کاری که پردازنده انجام میدهد» را با یک تشبیه ببینیم.
تو یک دستور پخت مینویسی (کد .java تو). این دستور به زبان انگلیسیِ آدمها است — خوانا برای تو، ولی سختافزار آن را نمیفهمد.
یک مترجم (به اسم کامپایلر، همان javac) میآید و دستور تو را به یک زبان بینالمللیِ استاندارد ترجمه میکند (چیزی شبیه اسپرانتو). به این زبانِ میانی میگویند بایتکد (فایل .class).
چرا این کار را میکند؟ چون این زبان میانی را همهٔ آشپزها میفهمند — فرقی نمیکند آشپز فرانسوی باشد (ویندوز) یا ژاپنی (لینوکس). هر کدام یک مترجمِ محلی به اسم JVM دارند که بایتکد را میگیرد و به زبان سختافزارِ خودش اجرا میکند.
پس مسیر کامل این است:
کد تو (.java) ──javac──▶ بایتکد (.class) ──JVM──▶ اجرا روی سختافزار واقعی
انگلیسیِ خوانا زبان میانی استاندارد زبان ماشینِ همان سیستمعامل
نکتهٔ ظریف: بایتکد مستقل از پلتفرم است — یعنی همان فایل .class روی هر سیستمی کار میکند. کاری که سیستمعاملها را متفاوت میکند، فقط آن آخرین قدم (JVM) است که برای هر سیستمعامل نسخهٔ جداگانه دارد.
JVM با بایتکد سه کار میکند:
- بارگذاری (load): فایل
.classرا پیدا و وارد حافظه میکند. - وارسی (verify): چک میکند بایتکد سالم و امن باشد.
- اجرا: ابتدا آن را خطبهخط تفسیر (interpret) میکند؛ و بخشهایی که خیلی تکرار میشوند («کد داغ») را در لحظهٔ اجرا به کد ماشینِ بومی کامپایل میکند تا سریع شوند. به این کامپایلِ حیناجرا میگویند JIT (Just-In-Time).
سه زیرسیستم اصلی JVM
- بارگذار کلاس (Class Loader) = بخش تدارکات و انبار: میرود فایلهای
.classرا پیدا میکند، میآورد داخل، و چک میکند سالم باشند. - نواحی دادهٔ زماناجرا (Runtime Data Areas) = فضای فیزیکی آشپزخانه: یخچال بزرگ (heap)، میز کارِ لحظهای (stack)، کتابخانهٔ دستورپختها (metaspace) — جایی که همهچیز حین کار زندگی میکند.
- موتور اجرا (Execution Engine) = سرآشپز و دستیارانش: کسی که واقعاً غذا را میپزد؛ شامل مفسر (interpreter)، کامپایلر هوشمند (JIT) و نظافتچی (Garbage Collector).
در ادامه هر سه دپارتمان را دقیق باز میکنیم.
مهمترین تغییرِ نگاه: «استاندارد» در برابر «محصول»
اینجا یک نکتهٔ سنیوری است که خیلیها اشتباه میفهمند:
دولت یک «قانون ساختمانسازی» مینویسد: «هر خانه باید در، پنجره، سقف و لولهکشی داشته باشد.» این قانون نمیگوید چطور بسازی، فقط میگوید چه خروجیای باید داشته باشد. این همان مشخصات (Specification) جاواست.
حالا چند شرکت پیمانکار میآیند و طبق این قانون خانه میسازند:
- Oracle با محصولی به اسم HotSpot (معروفترین و پیشفرض) — پر از تکنیکهای خفن برای سرعت و مدیریت حافظه.
- IBM با محصول OpenJ9 — تمرکز روی مصرف رم کمتر.
- Azul با محصول Zing — تمرکز روی مقیاسپذیریِ بسیار بالا.
نتیجهگیری این بخش، که برای مصاحبه طلاست:
JVM Specification = قانون و استاندارد (رفتار را تعریف میکند). HotSpot = محصولِ Oracle که آن قانون را پیاده کرده. تقریباً هر چیزی که در این فصل دربارهٔ GC، هدرِ آبجکت و JIT میگوییم، مربوط به رفتار خاصِ HotSpot است، نه الزامِ استاندارد. یک جونیور فکر میکند «جاوا = HotSpot». یک سنیور میداند جاوا یک استاندارد است و HotSpot فقط یکی از پیادهسازیهای آن — پس وقتی به مشکل میخورد، سراغِ تنظیماتِ خاصِ HotSpot میرود، نه اینکه فکر کند این ذاتِ زبان است.
بخش ۲ — بارگذاری کلاس: loading → linking → initialization
وقتی برنامهات اجرا میشود، JVM همهٔ کلاسها را یکجا لود نمیکند. فقط وقتی یک کلاس را لود میکند که برای اولین بار واقعاً به آن نیاز پیدا کنی. به این میگویند بارگذاری تنبل (lazy loading).
وقتی یک دایرةالمعارف ۱۰ جلدی میخری، همان لحظه هر ۱۰ جلد را نمیخوانی! فقط وقتی جلدِ «حرف ک» را باز میکنی که واقعاً کلمهای با «ک» لازم داشته باشی. JVM هم همینطور است: کلاس را دقیقاً لحظهای لود میکند که اولین بار به آن برسی (مثلاً وقتی new میزنی یا یک متد static صدا میکنی). این باعث میشود برنامه سریعتر بالا بیاید و حافظهٔ الکی اشغال نشود.
وقتی JVM بالاخره به یک کلاس نیاز پیدا کرد، آن را طبق استاندارد (JLS §12.4 و JVMS §5) در سه فاز با ترتیب دقیق آماده میکند. بیا با تشبیهِ استخدام یک کارمند جدید ببینیمشان.
فاز ۱: بارگذاری (Loading)
بخش «منابع انسانی» (همان ClassLoader) بایتهای فایل .class را (از روی دیسک، از داخل یک JAR، از شبکه، یا حتی بایتهای تولیدشده در لحظه) میخواند و یک پروندهٔ پرسنلی در حافظه میسازد — در جاوا این پرونده یک آبجکت از نوع Class<?> است که در heap قرار میگیرد.
هویتِ زماناجرای یک کلاس برابر است با ترکیبِ (نامِ کاملِ کلاس + همان ClassLoaderی که لودش کرده). یعنی اگر دقیقاً همان بایتها توسط دو ClassLoader مختلف لود شوند، JVM آنها را دو نوعِ کاملاً متفاوت میبیند!
بهخاطر همان قانون بالا، ممکن است خطای عجیبی ببینی: ClassCastException: com.x.Foo cannot be cast to com.x.Foo — یعنی «Foo را نمیتوان به Foo تبدیل کرد»! چطور ممکن است؟ چون این دو Foo را دو ClassLoader مختلف لود کردهاند و از نظر JVM دو نوعِ جدا هستند. این تله در سرورهای اپلیکیشن، OSGi و hot-reload خیلی رخ میدهد.
فاز ۲: لینک (Linking) — سه زیرمرحله
حالا که پرونده ساخته شد، باید کارمند را به سیستمِ شرکت وصل کنیم. سه زیرمرحله دارد:
وارسی (Verification) — بازرسی امنیتی: بایتکد از نظر امنیت و درستی چک میشود: آیا type-safe است؟ آیا عملیاتِ روی پشته معتبرند؟ آیا پرشِ غیرمجاز به جایی از کد ندارد؟ این ستون فقراتِ امنیتِ JVM است — همین وارسی است که نمیگذارد یک بایتکدِ دستکاریشده حافظه را خراب کند.
آمادهسازی (Preparation) — میز خالی: فیلدهای
staticساخته میشوند و مقدارِ پیشفرض میگیرند (عددها0، آبجکتهاnull، بولینهاfalse) — نه مقدار واقعی. مثل اینکه به کارمند یک میز خالی و دفترچهٔ سفید میدهی، نه ابزار کارِ واقعی.حل ارجاع (Resolution) — تبدیل وعده به آدرس: در کد نوشتهای «برو با کلاس
Databaseکار کن». این فعلاً یک نامِ نمادین (symbolic) و مبهم است. اینجا JVM آن نام را به ارجاعِ مستقیم (آدرسِ دقیقِ آن چیز در حافظه) تبدیل میکند. این مرحله هم میتواند تنبل باشد و تا لحظهٔ آخر عقب بیفتد.
فاز ۳: مقداردهی اولیه (Initialization)
حالا «روز اول کاری». JVM یک متدِ مخفیِ خاص به اسم <clinit> (مخففِ class initializer) را اجرا میکند. این متد را خودت نمینویسی؛ کامپایلر آن را از دو چیز میسازد و از بالا به پایین اجرا میکند:
- مقداردهیِ واقعیِ فیلدهای
static(همان0ها وnullهای مرحلهٔ قبل حالا مقدار واقعی میگیرند)، - و بلوکهای
static { ... }.
کِی اجرا میشود؟ فقط با «استفادهٔ فعال» از کلاس: با new، فراخوانی متد static، خواندن/نوشتنِ فیلد static (مگر یک ثابتِ زمانِ کامپایل باشد)، reflection، یا مقداردهیِ یک زیرکلاس.
JVM تضمین میکند <clinit> دقیقاً یکبار و بهصورتِ کاملاً thread-safe اجرا شود. یعنی اگر ۱۰۰ ترد همزمان اولین بار به این کلاس برسند، JVM خودش صف میبندد: فقط یکی مقداردهی میکند و بقیه منتظر میمانند تا آماده شود. بدون اینکه تو حتی یک کلمه synchronized بنویسی.
از همین تضمین، یک الگوی بسیار محبوب و هوشمند برای Singletonِ تنبل ساخته میشود:
public class Config {
private Config() {} // کسی از بیرون نمیتواند بسازدش
// این کلاس داخلی تا وقتی getInstance() صدا نخورد، اصلاً لود نمیشود (بارگذاری تنبل).
private static class Holder {
static final Config INSTANCE = new Config(); // <clinit> یکبار و thread-safe توسط JVM
}
public static Config getInstance() { return Holder.INSTANCE; }
}
تا وقتی کسی getInstance() را صدا نزند، کلاسِ Holder اصلاً وجودِ خارجی پیدا نمیکند (تنبل). لحظهای که صدا زده شد، JVM Holder را لود میکند و چون خودش تضمین میکند <clinit> فقط یکبار و امن اجرا شود، ما یک Singletonِ تنبل، امن در برابر چند-تردی، و بدونِ قفلِ دستی ساختیم — بهترینِ هر دو دنیا.
سلسلهمراتب ClassLoader و مدل واگذاری (Delegation)
بخشهای منابع انسانی (ClassLoaderها) بهصورت زنجیرهٔ والد-فرزند کار میکنند و یک قانونِ سفتوسخت دارند: «اول از رئیست بپرس.» به این میگویند مدل واگذاری به والد (parent-delegation): هر loader اول از والدش میخواهد کلاس را لود کند؛ فقط اگر والد نتوانست، خودش دستبهکار میشود.
Bootstrap ClassLoader (بومی/C++، هستهٔ جاوا را لود میکند: java.lang.* و ...)
└─ Platform ClassLoader (قبلاً «extension»؛ ماژولهای استانداردِ خودِ JDK)
└─ System/Application ClassLoader (classpathِ برنامهٔ تو)
└─ ClassLoaderهای سفارشی / وباپها (مثلاً Tomcat)
چرا این قانون وجود دارد؟ امنیت. بیا با یک مثال ببینیم:
فرض کن مدیرعامل (Bootstrap) مهرِ رسمی و اصلیِ شرکت را در گاوصندوقش دارد — معتبرترین مهرِ دنیا. یک هکر یک مهرِ تقلبیِ دقیقاً شبیهِ آن میسازد و یواشکی روی میزِ تو (پوشهٔ پروژهات) میگذارد؛ اسمش را هم میگذارد java.lang.String.
اگر قانونِ واگذاری نبود: تو نیاز به مهر داری، به میزِ خودت نگاه میکنی، مهرِ تقلبی را برمیداری و اسناد را با آن مهر میکنی. هکر برنده شد — کدِ Stringِ ویروسیِ او بهجای Stringِ امنِ جاوا اجرا میشود و کل سیستم آلوده میشود.
با قانونِ واگذاری: تو حق نداری سرخود از میزت مهر برداری. اول از مدیرعامل (Bootstrap) میپرسی. او میگوید «خودم نسخهٔ اصلیِ String را دارم، بیا این را استفاده کن.» پس مهرِ تقلبیِ روی میزت اصلاً دیده نمیشود. هکر شکست خورد.
واگذاری مانع میشود کسی با گذاشتنِ یک فایلِ همنام با کلاسهای هستهٔ جاوا (مثل java.lang.String) در پوشهٔ پروژه، نسخهٔ اصلی را سایه بزند (shadow کند). چون JVM همیشه اول از هسته و والدهای بالادست میپرسد، نسخهٔ اصلی و امن همیشه برنده است.
از Java 9 به بعد، loaderِ قدیمیِ «extension» به Platform ClassLoader تبدیل شد و کل ماجرا روی سیستمِ ماژول (JPMS) بنا شد. ضمناً چون Bootstrap با C++ نوشته شده و یک آبجکتِ جاوایی نیست، اگر SomeCoreClass.class.getClassLoader() را صدا بزنی، null میگیری (یعنی «من را Bootstrap لود کرده»).
در یک Tomcat ممکن است ۱۰ برنامهٔ وبِ مختلف همزمان اجرا شوند و هر کدام نسخهٔ متفاوتی از یک کتابخانه (مثلاً Log4j) بخواهند. اگر همه از والد بپرسند، تداخل میشود. برای همین Tomcat عمداً به loaderِ هر وباپ میگوید: «اول خودت پوشهات را بگرد؛ اگر پیدا نکردی، بعد از والد بپرس.» این معکوسِ واگذاریِ استاندارد است و هدفش ایزولهکردنِ برنامههاست تا همدیگر را خراب نکنند.
بخش ۳ — نواحی دادهٔ زماناجرا: حافظه کجا زندگی میکند؟
این بخش به تو نشان میدهد وقتی برنامه اجرا میشود، هر چیزی کجا ذخیره میشود. بیا با تشبیهِ آشپزخانه شروع کنیم و بعد جدولِ دقیق را ببینیم.
- Heap (هیپ) = یخچالِ بزرگ: هر چیزی که با
newمیسازی (همهٔ آبجکتها و آرایهها) اینجا با حجمِ زیاد نگهداری میشود. این ناحیه توسط Garbage Collector مدیریت میشود. - Stack (پشته) = میزِ کارِ لحظهایِ سرآشپز: چیزهای کوچک و موقتیِ هر ترد (متغیرهای محلی، ارجاعها، آدرسِ برگشت) اینجا میآیند و با تمامشدنِ متد پاک میشوند.
- Metaspace = کتابخانهٔ دستورپختها: اطلاعاتِ خودِ کلاسها (نه آبجکتها) اینجاست — ساختار، متدها، بایتکد.
حالا جدولِ کاملِ نواحی و اینکه هر کدام اگر پر شود چه خطایی میدهد:
| ناحیه | مشترک بین کیست؟ | چه چیزی نگه میدارد | خطای پرشدن |
|---|---|---|---|
| Heap | کلِ JVM | همهٔ آبجکتها و آرایهها | OutOfMemoryError: Java heap space |
| Metaspace | کلِ JVM | متادیتای کلاس، بایتکدِ متدها | OutOfMemoryError: Metaspace |
| JVM Stack | هر ترد جدا | فریمها (متغیرهای محلی، operandها، آدرس برگشت) | StackOverflowError |
| PC Register | هر ترد جدا | آدرسِ دستورِ بایتکدِ جاری | — |
| Native Method Stack | هر ترد جدا | فریمهای کدِ C/C++ (از طریق JNI) | — |
| Code Cache | کلِ JVM | کدِ ماشینِ کامپایلشده توسط JIT | CodeCache is full (JIT خاموش میشود) |
دو نکتهٔ مهم دربارهٔ این جدول:
۱. هر ترد پشتهٔ خودش را دارد. اگر یک متد بینهایت خودش را صدا بزند (بازگشتِ بیپایان)، پشته پر میشود و StackOverflowError میگیری. از طرف دیگر، اگر دهها هزار تردِ سیستمی (platform thread) بسازی، حافظهٔ بومیِ سیستم تمام میشود (هر پشته حدود ۵۱۲KB تا ۱MB است، قابلتنظیم با -Xss).
همین فشارِ «هر ترد یک پشتهٔ بومیِ گران میخواهد» است که virtual threadها حلش میکنند: آنها پشتهٔ بومی را بهصورت دائم اشغال (pin) نمیکنند، پس میتوانی میلیونها تایشان را داشته باشی. (فصلِ مخصوصِ خودش را دارد.)
۲. Metaspace جای PermGen را گرفت. این یکی از مهمترین تغییرات جاوا ۸ است و سؤالِ رایجِ مصاحبه — در بخش بعد کاملاً بازش میکنیم.
تمرکز روی Metaspace: چرا و چطور جای PermGen را گرفت
یک کارخانهٔ خودروسازی را تصور کن:
- Heap = پارکینگِ بزرگ: هر خودرویی که ساخته میشود (مثلاً ۱۰٬۰۰۰ خودروی یکمدل) اینجا پارک میشود. در جاوا، این خودروها همان آبجکتها هستند.
- Metaspace = اتاقِ نقشه و مهندسی: جایی که نقشهٔ ساختِ (blueprint) خودرو نگهداری میشود.
نکتهٔ کلیدی: برای ساختِ ۱۰٬۰۰۰ خودرو، فقط به یک نقشه نیاز داری، نه ۱۰٬۰۰۰ نقشه! در جاوا هم دقیقاً همین است — هزاران آبجکت از یک کلاس میسازی (همه در Heap)، اما اطلاعاتِ ساختاریِ خودِ کلاس فقط یک بار در Metaspace ذخیره میشود.
دقیقاً چه چیزی در Metaspace است؟ متادیتا (یعنی «داده دربارهٔ داده»): نامِ کلاس، لیستِ متدها و پارامترهایشان، لیستِ فیلدها و نوعِشان، اطلاعاتِ ارثبری، و Constant Pool. توجه: مقدارِ واقعیِ یک فیلد (مثلاً اینکه نامِ کاربر «علی» است) در Heap است؛ اما این واقعیت که «کلاسِ User یک فیلدِ name از نوع String دارد» در Metaspace است.
قبل از جاوا ۸ این فضا PermGen نام داشت و داخلِ Heap بود با یک سقفِ ثابت و کوچک. اگر برنامه کلاسهای زیادی لود میکرد، پر میشد و خطای معروفِ OutOfMemoryError: PermGen space میگرفتی. از جاوا ۸، این فضا به حافظهٔ بومیِ سیستمعامل (off-heap) منتقل شد و اسمش شد Metaspace. مزیت: حالا پویا رشد میکند و تا وقتی رمِ فیزیکی پر نشده، با یک سقفِ کوچکِ ثابت کرش نمیکند.
دو علتِ اصلی: (۱) نشتیِ ClassLoader — مخصوصاً در سرورهای وب: هر بار که کد را redeploy میکنی، Tomcat یک ClassLoaderِ جدید میسازد و همهٔ کلاسها را دوباره لود میکند؛ اگر ClassLoaderهای قدیمی درست دور ریخته نشوند، نسخههای تکراری در Metaspace انباشته میشوند. (۲) تولیدِ کلاسهای داینامیک — کتابخانههایی مثل Hibernate/Spring/CGLIB در زمان اجرا کلاس میسازند؛ اگر افسارگسیخته شوند، Metaspace را پر میکنند. برای همین در پروداکشن با -XX:MaxMetaspaceSize سقف میگذاری.
بخش ۴ — چیدمانِ آبجکت و هدرها (HotSpot، ۶۴ بیتی)
حالا که میدانیم آبجکتها در Heap زندگی میکنند، ببینیم یک آبجکت دقیقاً چه شکلی در حافظه است. هر آبجکتِ heap علاوه بر فیلدهای خودت، یک هدر (header) دارد که JVM برای مدیریتش لازم دارد:
[ mark word: 8 بایت ][ klass pointer: 4 بایت ][ فیلدهای تو... ][ padding تا مضربِ ۸ ]
- Mark word — چیزهای مدیریتی: hashCodeِ هویتی، بیتهای سن (age) برای GC، وضعیتِ قفل، و هنگام جابهجایی در GC اشارهگرِ forwarding.
- Klass pointer — به متادیتای کلاس در Metaspace اشاره میکند (یعنی «من از نوعِ کدام کلاسم؟»).
وقتی heap کوچکتر از ~۳۲GB است (پیشفرض)، HotSpot از compressed oops استفاده میکند: اشارهگرها را بهجای ۸ بایت در ۴ بایت ذخیره میکند (بهصورتِ آفستِ مقیاسدار). این یک صرفهجوییِ بزرگِ حافظه است — به همین برمیگردیم در بخشِ «چرا ۳۱GB بهتر از ۳۳GB است».
پس یک Objectِ خالی ۱۶ بایت میشود (۱۲ بایت هدر + ۴ بایت padding). آرایهها یک فیلدِ ۴ بایتیِ «طول» هم اضافه دارند.
همین سربارِ هدر توضیح میدهد چرا boxing گران است: یک Integerِ باکسشده حدود ۱۶ بایت (هدر + مقدار + padding) بهعلاوهٔ یک ارجاعِ جداگانه میگیرد، در حالی که یک intِ خام فقط ۴ بایت است. در حلقههای داغ و ساختماندادهٔ بزرگ، این تفاوت واقعاً حس میشود.
Java 24 با JEP 450 (هدرهای فشردهٔ آبجکت، فعلاً آزمایشی) هدر را به ۸ بایت میرساند. دانستنش خوب است، اما هنوز پیشفرض نیست.
بخش ۵ — مبانیِ Garbage Collection
اینجا میرسیم به جادوی جاوا: تو حافظه را دستی آزاد نمیکنی؛ یک زبالهجمعکن (Garbage Collector) خودکار این کار را میکند. اما چطور میفهمد کدام آبجکت «زباله» است؟
GC بر پایهٔ «قابلیت دسترسی» کار میکند، نه شمارش
تصور کن هر آبجکت یک بادکنک است و ارجاعها نخهایی که بادکنکها را به هم و به یک لنگرِ ثابت (GC root) وصل میکنند. لنگرها چیزهایی هستند که همیشه «زنده»اند: متغیرهای محلیِ تردهای درحالاجرا، فیلدهای static، و ارجاعهای JNI.
GC از لنگرها شروع میکند و هر بادکنکی را که با نخها بتوان به یک لنگر رسید، «زنده» علامت میزند (mark). هر بادکنکی که هیچ مسیری به لنگر نداشته باشد، زباله است و آزاد میشود.
چون چرخهها را درست مدیریت میکند. فرض کن آبجکت A به B اشاره کند و B به A، ولی هیچکدام به یک لنگر وصل نباشند — یک «جزیرهٔ جدا». در روشِ سادهٔ شمارشِ ارجاع، شمارندهٔ هر دو ۱ است پس هرگز آزاد نمیشوند (نشتی!). اما GCِ ردیابیمحورِ (tracing) جاوا چون از لنگر شروع میکند و به این جزیره نمیرسد، هر دو را درست تشخیصِ مرده میدهد.
فرضیهٔ نسلی (Generational Hypothesis)
یک مشاهدهٔ تجربی که کلِ طراحیِ GCهای مدرن رویش بنا شده: بیشترِ آبجکتها جوان میمیرند. یعنی اکثرِ آبجکتها خیلی زود پس از ساختهشدن بیمصرف میشوند (مثل متغیرهای موقتِ داخلِ یک متد). GC از این حقیقت با تقسیمِ heap بهره میبرد:
نسلِ جوان (Young): [ Eden | Survivor S0 | Survivor S1 ] نسلِ قدیم (Old / Tenured)
- آبجکتهای تازه در Eden ساخته میشوند (تخصیصِ فوقسریع، فقط با جابهجاییِ یک اشارهگر).
- یک GC جوان (minor GC) آبجکتهای زندهٔ Eden و یک Survivor را به Survivor دیگر کپی میکند و سنشان را یکی زیاد میکند. چون فقط آبجکتهای زنده را لمس میکند و بیشترشان مردهاند، خیلی ارزان است.
- آبجکتی که بهاندازهٔ کافی چرخهٔ جوان را رد کند (
-XX:MaxTenuringThreshold، تا ۱۵)، به نسلِ قدیم (Old) «ترفیع (promote/tenure)» مییابد. - یک GC قدیم (major GC) نسلِ قدیم را جمع میکند؛ و یک Full GC همهچیز را (جوان + قدیم + اغلب metaspace). Full GC همان موردِ گرانقیمتی است که باید ازش پرهیز کنی.
برای اینکه GC بتواند بیخطر آبجکتها را جابهجا کند و لنگرها را بشمارد، JVM لحظهای همهٔ تردهای برنامه را در یک نقطهٔ امن (safepoint) متوقف میکند — به این توقف میگویند Stop-the-World (STW). هر GCی مقداری STW دارد. GCهای مدرن این توقف را با انجامِ بیشترِ کار بهصورتِ همزمان (concurrent) با برنامه کمینه میکنند. سیستمهای حساس به تأخیر، زندگی و مرگشان با طول و تعدادِ همین توقفهاست.
ارجاعهای ضعیف: weak / soft / phantom
جاوا چند نوع ارجاعِ خاص دارد که به تو اجازه میدهند به GC بگویی «این آبجکت آنقدرها هم مهم نیست»:
SoftReference— فقط زیرِ فشارِ حافظه پاک میشود. مناسبِ کش (ولی مراقب باش، میتواند نشتی بدهد).WeakReference— در اولین GC بعد از اینکه فقط ضعیف قابلدسترس شد پاک میشود. مبنایWeakHashMap.PhantomReference— برای پاکسازیِ قطعیِ منابعِ بومی بعد از مرگِ آبجکت (کلاسِCleaner، جایگزینِ مدرنِfinalize()).
بخش ۶ — چشمانداز GCها: کدام کِی؟
جاوا چند Garbage Collector مختلف دارد و میتوانی با یک فلگ انتخابشان کنی. هر کدام یک مصالحه (trade-off) بین throughput (کارِ کلیِ انجامشده) و latency (کوتاهیِ توقفها) دارند.
| Collector | فلگ | مدلِ توقف | بهترین برای | مصالحه |
|---|---|---|---|---|
| Serial | -XX:+UseSerialGC |
کاملاً STW، تکرشتهای | heapِ کوچک، کانتینرِ تکCPU، ابزارِ CLI | سادهترین و کمسربارترین؛ ولی در مقیاسِ بزرگ توقفهای طولانی |
| Parallel | -XX:+UseParallelGC |
کاملاً STW، چندرشتهای | jobهای batch، بیشینهکردنِ throughput | بهترین throughput؛ بدترین tail-latency |
| G1 | -XX:+UseG1GC (پیشفرض از Java 9) |
مارکِ عمدتاً concurrent، تخلیهٔ STW | همهمنظوره، heapِ بزرگ، تعادلِ latency/throughput | هدفِ زمانِ توقف میدهی (-XX:MaxGCPauseMillis، پیشفرض ۲۰۰ms) |
| ZGC | -XX:+UseZGC |
concurrent، توقفِ زیرِ ۱ms | heapِ بسیار بزرگ (تا ترابایت)، SLAی سختِ تأخیر | throughputِ کمی کمتر، سربارِ CPU/حافظهٔ بیشتر |
| Shenandoah | -XX:+UseShenandoahGC |
concurrent، توقفِ زیرِ ۱ms | تأخیرِ پایین، بیلدهای Red Hat | مشابهِ ZGC |
در همهٔ JDKهای مدرن — از جمله Java 17 و Java 21 — Collectorِ پیشفرض همچنان G1 است. ZGC و Shenandoah اختیاری (opt-in) هستند. اگر کسی بپرسد «پیشفرضِ ۱۷ و ۲۱؟»، جوابِ درست: G1 (که از Java 9 جایگزینِ Parallel شد).
G1 با کمی جزئیات
G1 (مخففِ Garbage-First) هیپ را به حدودِ ۲۰۴۸ ناحیهٔ (region) مساوی تقسیم میکند (هر کدام ۱ تا ۳۲MB). هر ناحیه بهصورتِ پویا برچسبِ Eden، Survivor، Old یا Humongous (آبجکتهای خیلی بزرگتر از نصفِ یک ناحیه) میگیرد. G1 عمدتاً بهصورتِ concurrent مارک میکند، بعد توقفهای تخلیه (evacuation) ای انجام میدهد که آبجکتهای زنده را اول از ناحیههایی بیرون میکشد که بیشترین زباله را دارند — برای همین اسمش «زبالهاول» است. تو یک هدفِ توقف میدهی، نه یک چیدمانِ دقیق؛ G1 خودش تصمیم میگیرد چند ناحیه را جمع کند تا به هدفت برسد. G1 عملاً Full GCهای چندثانیهایِ collectorِ قدیمیِ CMS را حذف کرد (CMS در Java 14 کاملاً برداشته شد).
ZGC و Generational ZGC — نسخهها را دقیق بگو
ZGC یک collectorِ concurrent، ناحیهمحور و فشردهساز است که با ترفندهایی به اسمِ اشارهگرهای رنگی (colored pointers) و load barrier آبجکتها را حین اجرای برنامه جابهجا میکند. تیترِ اصلیاش: توقفها زیرِ یک میلیثانیهاند (معمولاً ۰٫۰۵ تا ۰٫۵ms) و با بزرگشدنِ heap رشد نمیکنند — تا heapهای چند-ترابایتی مقیاس میپذیرد.
- Java 15: ZGC آمادهٔ تولید شد (JEP 377).
- Java 21: نسخهٔ Generational ZGC اضافه شد (JEP 439)، ولی اختیاری بود:
-XX:+UseZGC -XX:+ZGenerational. صرفِ-XX:+UseZGCنسخهٔ قدیمیِ غیرنسلی را میداد. - Java 23: حالتِ نسلی برای ZGC پیشفرض شد (JEP 474) و فلگِ
ZGenerationalمنسوخ اعلام شد. - Java 24: حالتِ غیرنسلی کاملاً حذف شد (JEP 490).
دقیقگفتنِ این نسخهها به مصاحبهگر نشان میدهد پلتفرم را واقعاً دنبال میکنی.
وقتی heap بزرگ است و نیازِ سختی به کوتاهیِ tail-latency داری (سیستمهای ترید، سرویسهای فوقکمتأخیر) که حتی توقفهای ~۱۰۰-۲۰۰ میلیثانیهایِ G1 هم برایت زیاد است. برای سرویسهای عمومی روی G1 بمان — معمولاً throughputِ بهتر و سربارِ حافظهٔ کمتری دارد و پیشفرضِ آزموده است.
بخش ۷ — کامپایلِ JIT، tiered و inlining
یادت هست گفتیم JVM اول کد را تفسیر میکند و بخشهای داغ را کامپایل؟ حالا دقیقترش میکنیم.
HotSpot ابتدا بایتکد را تفسیر (interpret) میکند و همزمان از آن پروفایل میگیرد: هر متد چند بار صدا خورده؟ کدام branchها بیشتر گرفته میشوند؟ کدام تایپها واقعاً میآیند؟ متدهای داغ توسط دو کامپایلر به کدِ ماشین تبدیل میشوند:
- C1 (client): کامپایلِ سریع، بهینهسازیِ سبک → استارتاپِ سریع.
- C2 (server): کامپایلِ کند، بهینهسازیِ تهاجمی (inlining، بازکردنِ حلقه، حذفِ کدِ مرده) → اوجِ کارایی.
مفسر مثل یک آشپزِ مبتدی است: بایتکد را خطبهخط میخواند و اجرا میکند (سرعتِ معمولی، ولی همین الان شروع میکند). JIT مثل سرآشپزی است که حواسش هست کدام غذا صد بار سفارش داده شده؛ آن دستور را از قبل به «حرکاتِ عضلانیِ خودکار» (کدِ ماشینِ بومی) تبدیل میکند تا دفعهٔ بعد با سرعتِ نور بپزد.
Tiered compilation (پیشفرض) هر دو کامپایلر را در ۵ سطح ترکیب میکند:
Level 0: مفسر (interpreter)
Level 1: C1، بدونِ پروفایل (متدهای بدیهی)
Level 2: C1، پروفایلِ محدود
Level 3: C1، پروفایلِ کامل ← بیشترِ متدها اینجا گرم میشوند
Level 4: C2، کاملاً بهینه ← داغترین متدها به اینجا میرسند
کد از تفسیر شروع میشود، بعد از گرمشدن با پروفایل توسط C1 کامپایل میشود، و مسیرهای واقعاً داغ به C2 «فارغالتحصیل» میشوند. کدِ کامپایلشده در code cache زندگی میکند.
اگر code cache پر شود، JIT خاموش میشود و همهچیز به مفسرِ کُند برمیگردد — یعنی افتِ ناگهانیِ کارایی. لاگش CodeCache is full است. در برنامههای خیلی بزرگ گاهی باید سقفش را زیاد کنی (-XX:ReservedCodeCacheSize).
Inlining — باارزشترین بهینهسازی
Inlining یعنی جایگزینکردنِ فراخوانیِ یک متد با بدنهٔ خودِ آن متد. چرا مهم است؟ چون وقتی بدنه سرِ جایش کپی شد، همهٔ بهینهسازیهای دیگر میتوانند از مرزِ متد عبور کنند و با هم ترکیب شوند. C2 متدهای کوچک و داغ را تهاجمی inline میکند.
خیلیها میترسند getter/setter اضافه کنند چون فکر میکنند «هزینهٔ فراخوانیِ متد» دارد. در عمل C2 این متدهای کوچک را inline میکند و کاملاً محو میشوند — انگار مستقیم به فیلد دست زده باشی. پس نگرانِ کاراییِ getterها نباش.
تحلیلِ فرار (Escape Analysis) و جایگزینیِ اسکالر
C2 سعی میکند اثبات کند آیا یک آبجکت از متد/تردِ سازندهاش فرار میکند (یعنی جایی بیرون هم دیده میشود) یا نه. اگر ثابت شود هرگز فرار نمیکند:
- جایگزینیِ اسکالر (scalar replacement): آبجکت اصلاً روی heap ساخته نمیشود؛ فیلدهایش مستقیم در رجیستر/پشته مینشینند. برای همین یک حلقهٔ داغ که یک
Pointِ موقت میسازد میتواند صفر زباله تولید کند. - حذفِ قفل (lock elision): اگر روی آبجکتی که فرار نمیکند
synchronizedباشد، آن قفل کلاً حذف میشود.
// C2 میتواند اثبات کند p فرار نمیکند؛ پس آبجکتِ موقت شاید هرگز به heap نرسد.
int sumOfSquares(int[] xs, int[] ys) {
int total = 0;
for (int i = 0; i < xs.length; i++) {
Point p = new Point(xs[i], ys[i]); // نامزدِ جایگزینیِ اسکالر
total += p.x * p.x + p.y * p.y;
}
return total;
}
تحلیلِ فرار تضمین نیست — بهترینتلاش است و میتواند شکست بخورد (مثلاً وقتی آبجکت به یک متدِ inlineنشده پاس شود). درستیِ برنامهات را بر پایهٔ آن طراحی نکن. و اگر داری با JMH بنچمارک میگیری، بدان که به همین دلیل JMH از Blackhole استفاده میکند تا نگذارد کامپایلر کدِ «بهظاهر بیمصرفِ» تو را کلاً حذف کند.
دیاپتیمایزیشن (Deoptimization)
C2 شرطبندیهای حدسی میکند (مثلاً «این نقطهٔ فراخوانی همیشه فقط ArrayList دیده، پس فرض میکنم همیشه همین است و کدِ بهینه میسازم»). اگر یک روز واقعیت این فرض را نقض کند (ناگهان یک LinkedList بیاید)، JVM آن کدِ کامپایلشده را دور میریزد و موقتاً به مفسر برمیگردد و شاید دوباره کامپایل کند. این بهصورتِ یک افتِ کاراییِ کوتاه بعد از ظهورِ یک تایپِ جدید دیده میشود.
بخش ۸ — فلگهای کلیدی که هر سنیور باید بداند
# اندازهٔ heap — در پروداکشن min == max بگذار تا از توقفِ resize و تکهتکهشدن جلوگیری شود
-Xms4g -Xmx4g
# کانتینر-آگاه (از Java 10+ پیشفرض فعال): heap را درصدی از حافظهٔ کانتینر بگیر
-XX:MaxRAMPercentage=75.0
# انتخابِ collector
-XX:+UseG1GC # پیشفرض
-XX:+UseZGC # از Java 23+ بهصورتِ نسلی
-XX:MaxGCPauseMillis=100
# اندازهٔ پشتهٔ هر ترد
-Xss512k
# سقفِ Metaspace (وگرنه تا اتمامِ حافظهٔ بومی رشد میکند)
-XX:MaxMetaspaceSize=256m
# هنگام OOM: heap dump بگیر و سریع بمیر (برای عیبیابیِ بعدی)
-XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/dumps
-XX:+ExitOnOutOfMemoryError
# لاگِ GC (یکپارچه، از Java 9+)
-Xlog:gc*:file=gc.log:time,uptime,level,tags:filecount=5,filesize=20m
اگر min و max یکی نباشند، heapِ متعهدشده مدام کوچک و بزرگ میشود که خودش باعثِ Full GC و page fault میگردد. با ثابتکردنِ اندازه، این نوسان حذف میشود. در کانتینرها MaxRAMPercentage را ترجیح بده تا JVM به محدودیتِ cgroup احترام بگذارد، نه اینکه کلِ رمِ هاست را ببیند و بعد OOM-kill شود.
بخش ۹ — نشتیِ حافظه و طبقهبندیِ OutOfMemoryError
در زبانی که GC دارد، نشتیِ حافظه یعنی قابلیتِ دسترسیِ ناخواسته: آبجکتهایی که دیگر کارت با آنها تمام شده، ولی هنوز از یک لنگرِ GC قابلدسترساند، پس GC حق ندارد آزادشان کند و روی هم انباشته میشوند.
منابعِ کلاسیکِ نشتی:
- کالکشنهای
staticکه فقط رشد میکنند (static Map cache = ...که هرگز چیزی از آن حذف نمیشود). - کشهای بیکران — بهجایش از سقفِ اندازه استفاده کن (Caffeine، LRU).
- listenerها/callbackهایی که هرگز unregister نمیشوند — سوژه، observer را برای همیشه نگه میدارد.
ThreadLocalدر thread poolها — تردِ poolشده هرگز نمیمیرد، پس مقدارِThreadLocalهم هرگز آزاد نمیشود؛ همیشه درfinallyآن راremove()کن.- نشتیِ ClassLoader — یک ارجاعِ باقیمانده به کلاسهای وباپ، کلِ ClassLoader (و همهٔ کلاسهایش در Metaspace) را هنگامِ redeploy زنده نگه میدارد.
انواعِ OutOfMemoryError (هر کدام معنای متفاوتی دارد)
| پیام | معنا | علتِ معمول |
|---|---|---|
Java heap space |
heap پر شد و GC نتوانست کافی آزاد کند | نشتی یا heapِ کوچک |
GC overhead limit exceeded |
بیش از ۹۸٪ زمان صرفِ GC میشود، کمتر از ۲٪ آزاد میشود | heap تقریباً پر، thrashing |
Metaspace |
فضای متادیتای کلاس تمام شد | نشتیِ classloader، کلاسهای داینامیکِ زیاد |
unable to create new native thread |
حافظهٔ بومیِ OS یا محدودیتِ ulimit پر شد | نشتیِ ترد، -Xssِ بزرگ × تردهای زیاد |
Direct buffer memory |
ByteBufferِ مستقیمِ off-heap تمام شد |
بافرهای مستقیمِ نشتی |
Requested array size exceeds VM limit |
آرایهای نزدیکِ Integer.MAX_VALUE خواستی |
باگِ منطقی |
این یک Error است، نه Exception. گرفتنش تقریباً همیشه اشتباه است، چون JVM ممکن است در وضعیتِ غیرقابلِبازیابی باشد. بگذار برنامه سریع بمیرد و با heap dump عیبش را پیدا کن.
بخش ۱۰ — چطور یک GC log را بخوانیم
JVMهای مدرن (۹+) از لاگِ یکپارچه (-Xlog:gc*) استفاده میکنند. یک توقفِ جوانِ G1 این شکلی است:
[2.335s][info][gc,start] GC(12) Pause Young (Normal) (G1 Evacuation Pause)
[2.340s][info][gc ] GC(12) Pause Young (Normal) (G1 Evacuation Pause) 512M->48M(1024M) 5.219ms
اینطور بخوانش: GC شمارهٔ ۱۲ یک توقفِ تخلیهٔ جوان بود؛ heap از ۵۱۲M قبل ← ۴۸M بعد رفت (یعنی ۴۶۴M آزاد شد)، کلِ heap ۱۰۲۴M است، و توقف ۵٫۲۱۹ms طول کشید.
دنبالِ این نشانهها بگرد:
- افزایشِ تعداد و طولِ توقفها = دردسر در راه است.
- دیدنِ
Pause Fullروی G1/ZGC = پرچمِ قرمز؛ بررسی کن (جهشِ تخصیص، آبجکتهای humongous، heapِ کوچک). - بالارفتنِ پیوستهٔ اندازهٔ زنده پس از هر Full GC = نشتی (هر جمعآوری کمتر آزاد میکند؛ کف مدام بالا میرود).
to-space exhausted= آبجکتهای جوان خیلی سریع ترفیع مییابند؛ اندازهها را تیون کن.
برای دیدنِ محتوای heap یک dump بگیر (-XX:+HeapDumpOnOutOfMemoryError یا jmap -dump) و در Eclipse MAT بازش کن — «dominator tree» و «leak suspects»ِ آن دقیقاً میگویند چه چیزی حافظه را نگه داشته. برای تشخیصِ زنده: jstat -gcutil <pid> 1s، jcmd <pid> GC.heap_info و JFR (-XX:StartFlightRecording).
بخش ۱۱ — دامها و نکاتِ ظریفِ رایج
- صداکردنِ
System.gc()— یک درخواست است نه دستور؛ معمولاً یک Full GCِ گران راه میاندازد و آسیب میزند. با-XX:+DisableExplicitGCخنثیاش کن. finalize()— منسوخ، غیرقابلِپیشبینی، میتواند آبجکت را دوباره زنده کند و جمعآوری را عقب بیندازد. ازCleanerیا try-with-resources استفاده کن.- object pooling برای آبجکتهای کوچک — تخصیص در نسلِ جوان تقریباً رایگان است؛ poolکردن معمولاً با نگهداشتنِ آبجکتها تا نسلِ قدیم فشارِ GC را بیشتر میکند. فقط منابعِ واقعاً گران (کانکشن، ترد) را pool کن.
-Xmxِ عظیم — heapِ بزرگتر یعنی توقفهای طولانیتر و Full GCهای دیرتر ولی بدتر. بزرگتر همیشه بهتر نیست.
بالای ~۳۲GB، JVM دیگر compressed oops را نمیتواند استفاده کند، پس هر ارجاع از ۴ به ۸ بایت دوبرابر میشود. این سربار آنقدر زیاد است که یک heapِ ۳۱GB میتواند آبجکتهای قابلاستفادهٔ بیشتری از یک heapِ ۳۳GB نگه دارد! پس یا زیرِ ۳۲GB بمان، یا اگر رد میشوی، بهقدرِ کافی رد شو که ارزشش را داشته باشد.
بخش ۱۲ — سؤالاتِ مصاحبه (با جواب)
در هر دو G1. از Java 9 پیشفرض بوده (جایگزینِ Parallel شد). ZGC و Shenandoah اختیاریاند.
تا Java 20 تکنسلی بود. Generational ZGC در Java 21 آمد (JEP 439) ولی اختیاری با -XX:+UseZGC -XX:+ZGenerational. در Java 23 برای ZGC پیشفرض شد (JEP 474) و حالتِ غیرنسلی در Java 24 حذف شد (JEP 490).
PermGen (پیش از Java 8) متادیتای کلاس را در ناحیهای با اندازهٔ ثابت از heap نگه میداشت و باعثِ OutOfMemoryError: PermGen spaceِ مکرر در redeployها میشد. Java 8 آن را با Metaspace در حافظهٔ بومی که پویا رشد میکند جایگزین کرد. هنوز میشود نشتش داد (نشتیِ classloader)؛ با -XX:MaxMetaspaceSize محدودش کن.
بیشترِ آبجکتها جوان میمیرند. پس heap تقسیم میشود و young GC از یک collectorِ کپیکننده استفاده میکند که فقط آبجکتهای زنده را لمس میکند (بازماندگان را کپی و Eden را یکجا reset میکند). چون بیشترِ آبجکتهای جوان مردهاند، کارِ کمی میکند و رایگان فشردهسازی هم میشود — بدونِ fragmentation، با هزینهای متناسبِ بازماندگان نه زباله.
safepoint نقطهای است که وضعیتِ ترد سازگار و برای JVM شناختهشده است تا بتوان بیخطر متوقفش کرد. حتی ZGC هم به توقفهای کوتاهِ STW نیاز دارد (مثلاً شروع/پایانِ مارکِ ریشهها) تا یک snapshotِ سازگار بسازد؛ اما بخشِ عمدهٔ مارک/جابهجایی concurrent است. «concurrent» یعنی بیشترِ کار با برنامه همپوشانی دارد، نه صفر توقف.
با اشارهگرهای رنگی (بیتهای متادیتا داخلِ خودِ اشارهگر) بهعلاوهٔ load barrier: وقتی برنامه یک ارجاع را load میکند، barrier آن را چک/تصحیح میکند و به ZGC اجازه میدهد آبجکتها را همزمان با برنامه جابهجا کند. کارِ توقف متناسبِ تعدادِ ریشههای GC است (کراندار)، نه اندازهٔ heap — پس با رشدِ heap تا ترابایت، توقفها مسطح میمانند.
C2 تحلیل میکند آیا آبجکت از متد/ترد فرار میکند. اگر نه: جایگزینیِ اسکالر (آبجکت روی heap ساخته نمیشود؛ فیلدها رجیستر/پشته میشوند ← صفر زباله) و حذفِ قفل (حذفِ synchronizedِ بیفایده). بهترینتلاش است، نه تضمین.
مفسر اول با پروفایلگیری اجرا میکند. C1 (سریع، بهینهسازیِ سبک) متدهای گرم را با پروفایلِ کامل کامپایل میکند (سطح ۳)؛ داغترینها به C2 (کند، بهینهسازیِ تهاجمی → سطح ۴) میرسند. این استارتاپِ سریع (C1) را با اوجِ throughput (C2) متعادل میکند.
class Cache {
private static final Map<Key, Value> M = new HashMap<>();
static void put(Key k, Value v) { M.put(k, v); } // هرگز evict نمیشود
}
یک Mapِ static که فقط رشد میکند برای همیشه از یک لنگرِ GC قابلدسترس است ← رشدِ بیکرانِ heap ← OutOfMemoryError: Java heap space. اصلاح: محدودش کن (Caffeine/LRU) یا اگر باید خودکار جمع شوند از WeakHashMap/SoftReference استفاده کن.
Integer a = 127, b = 127;
Integer c = 128, d = 128;
System.out.println((a == b) + " " + (c == d));
چاپ میکند: true false. متدِ Integer.valueOf بازهٔ −۱۲۸ تا ۱۲۷ را کش میکند، پس a و b همان آبجکتِ کششدهاند (== درست است)، اما 128 خارجِ کش است ← دو آبجکتِ جدا (== نادرست). این autoboxing + کشِ Integer است و دقیقاً برای همین برای تایپهای باکسشده همیشه از .equals() استفاده میکنی.
هویتِ زماناجرای کلاس برابرِ (نام، ClassLoaderِ تعریفکننده) است. اگر دو ClassLoader هر کدام com.x.Foo را لود کنند، JVM آنها را دو نوعِ متفاوت میبیند و castشان ClassCastException میدهد، حتی اگر سورس یکی باشد. رایج در app serverها، OSGi و hot-reload.
امضای کلاسیکِ نشتیِ حافظه: هر Full GC کمتر آزاد میکند و کفِ زنده بالا میرود. heap dump بگیر، در Eclipse MAT باز کن، با dominator tree / leak suspects لنگرِ نگهدارنده را پیدا کن — اغلب یک کالکشنِ static، کشِ بیکران، یا ThreadLocalی که در pool مقدارش remove() نشده.
بالای ~۳۲GB، JVM compressed oops را خاموش میکند، پس هر ارجاع از ۴ به ۸ بایت رشد میکند. این سربار میتواند ظرفیتِ قابلاستفاده را چنان کم کند که heapِ ۳۱GB آبجکتهای زندهٔ بیشتری از ۳۳GB نگه دارد. فقط وقتی واقعاً لازم است از مرز عبور کن.
نه — یک اشاره است که JVM میتواند نادیده بگیرد، و -XX:+DisableExplicitGC آن را بیاثر میکند. وقتی هم اجرا شود معمولاً یک Full STW GCِ گران را اجبار میکند، پس تکیه بر آن در کدِ اپلیکیشن ضدالگوست.
متغیرهای primitiveِ محلی و ارجاعهای آبجکت روی پشتهٔ ترد (در فریم) هستند؛ خودِ آبجکتها روی heap. متادیتای کلاس و فیلدهای static در Metaspace (بومی) هستند، ولی آبجکتی که فیلدِ static به آن اشاره میکند همچنان روی heap است. خلاصهٔ تیز: «primitiveها و ارجاعها روی stack، آبجکتها روی heap» — با این نکته که تحلیلِ فرار میتواند بعضی آبجکتهای کوتاهعمر را کلاً خارج از heap نگه دارد.
نکاتِ سنیور و موارد پیشرفته
تا اینجا فهمیدیم حافظه کجا زندگی میکند و GC چطور تمیز میکند. اما در پروداکشن، چیزی که تو را ساعت سه صبح بیدار میکند معمولاً همین «لایهٔ پنهانِ» زیرِ آن مفاهیم است: چرا کانتینرت با اینکه -Xmx رعایت شده OOM-kill میشود، چرا یک ترد بیگناه کلِ JVM را قفل میکند، چرا کدی که «درست» است روی چند هسته نتیجهٔ کهنه میبیند. این بخش دقیقاً همان لایه است.
(۱) مدلِ حافظهٔ جاوا (JMM) — happens-before، volatile و انتشارِ امن؛ (۲) RSS در برابر heap و اینکه چرا کانتینر kill میشود؛ (۳) TLAB و مسیرِ واقعیِ تخصیص؛ (۴) card table و remembered set و ارجاعهای بیننسلی؛ (۵) time-to-safepoint و دامِ حلقهٔ شمارشی؛ (۶) حالتهای قفل در mark word؛ (۷) compressed class space، آبجکتهای humongous، inline cache؛ (۸) startup و warmup (AppCDS/AOT/CRaC/GraalVM)؛ و ۸ سؤالِ سختِ سنیوری در انتها.
۱) مدلِ حافظهٔ جاوا (JMM): چیزی که فصل اصلی جا انداخت
فصل دربارهٔ «حافظه» بود، اما یک تکهٔ حیاتیاش کجاست؟ مدلِ حافظه. Heap بین همهٔ تردها مشترک است، ولی هر هسته cache و بافرِ نوشتنِ خودش را دارد و کامپایلر/CPU مجازند دستورها را جابهجا کنند. پس این سؤال اصلاً بدیهی نیست: «اگر ترد A فیلدی را بنویسد، ترد B کِی آن را میبیند؟»
A روی دفترچهٔ شخصیاش (cacheِ هسته) مینویسد و فکر میکند B هم دیده. ولی تا وقتی روی وایتبردِ مشترک (حافظهٔ اصلی) کپی نکند، B نسخهٔ کهنه را میبیند. JMM قوانینِ «کِی باید روی وایتبرد کپی کنی» است.
قلبِ JMM رابطهٔ happens-before است: اگر عملِ X، happens-before عملِ Y باشد، اثرِ X تضمیناً برای Y دیده میشود. مهمترین یالها:
- قفلِ مانیتور:
unlockروی یک قفل، happens-before هرlockبعدیِ همان قفل. - volatile: نوشتنِ یک فیلدِ
volatile، happens-before هر خواندنِ بعدیِ همان فیلد. - شروع/پیوستنِ ترد:
thread.start()happens-before کدِ داخلِ آن ترد؛ و کدِ ترد happens-before بازگشتِthread.join(). - تعدی (transitivity): اگر A→B و B→C پس A→C.
میدهد: دیدهشدن (visibility)، جلوگیری از جابهجاییِ دستور (ordering)، و خواندن/نوشتنِ اتمیکِ long/double. نمیدهد: اتمیکبودنِ عملیاتِ مرکب. count++ روی فیلدِ volatile هم مسابقهای است، چون read-modify-write است. برای شمارنده از AtomicInteger/LongAdder یا VarHandle (CAS) استفاده کن.
دسترسیِ همزمانِ بدونِ همگامسازی که حداقل یک نوشتن دارد = data race و رفتارش تعریفنشده است؛ نه فقط ممکن است مقدار کهنه ببینی، بلکه ممکن است اصلاً هرگز آپدیت را نبینی (کامپایلر مقدار را در رجیستر hoist کند و حلقهات ابدی شود). این باگ روی لپتاپ x86 «کار میکند» و روی سرورِ ARM یا زیرِ بارِ سنگین میترکد — بدترین نوعِ باگ.
انتشارِ امن (safe publication) و فیلدهای final. چطور یک آبجکت را بیقفل به تردِ دیگر بدهیم؟ JMM تضمین میکند اگر آبجکت درست ساخته شده باشد (یعنی this در حینِ سازنده فرار نکند)، فیلدهای final آن بلافاصله پس از پایانِ سازنده برای همهٔ تردها دیده میشوند — این «انجمادِ فیلدِ final» دقیقاً همان چیزی است که String و آبجکتهای immutable را امن میکند: میتوانی بیهیچ همگامسازیای بینِ تردها پاسشان بدهی. راههای امنِ دیگر: نوشتن در فیلدِ volatile/AtomicReference، استفاده از static initializer (همان تضمینِ <clinit>)، یا عبور از یک collectionِ concurrent.
// double-checked locking درست: instance حتماً باید volatile باشد،
// وگرنه تردِ دیگر میتواند ارجاعِ غیرnull ولی نیمهساخته ببیند.
class Lazy {
private static volatile Lazy instance; // بدونِ volatile شکسته است
static Lazy get() {
Lazy r = instance;
if (r == null) synchronized (Lazy.class) {
r = instance;
if (r == null) instance = r = new Lazy();
}
return r;
}
}
در عمل کمتر خودت DCL بنویس؛ الگوی Holder (که فصل نشان داد) سادهتر و بیخطاتر است. DCL را بدان تا در code review دام volatileِ فراموششده را بگیری.
۲) RSS در برابر heap: چرا کانتینر با -Xmx2g در ۳GB kill میشود
شایعترین «معمای پروداکشن»: heapِ زنده ۱.۲GB است، -Xmx2g گذاشتهای، اما کانتینر روی ۳GB به OOMKilled (exit 137) میخورد. چرا؟ چون heap فقط یک تکه از حافظهٔ کلِ پروسه است. kernel بر اساسِ RSS (حافظهٔ فیزیکیِ کلِ پروسه) میکُشد، نه بر اساسِ heapِ جاوا.
یکخط توضیح از تجزیهٔ فضای پروسه؛ RSS جمعِ همهٔ اینهاست، نه فقط heap:
flowchart TD
RSS["Container RSS (what the kernel OOM-kills on)"] --> Heap["Java Heap (-Xmx)"]
RSS --> Meta["Metaspace + Compressed Class Space"]
RSS --> Code["JIT Code Cache"]
RSS --> Stacks["Thread stacks (N x -Xss)"]
RSS --> Direct["Direct / mapped ByteBuffers (NIO, Netty)"]
RSS --> GC["GC structures (card table, RSets, mark bitmaps)"]
RSS --> Native["Native libs, JNI, malloc arenas"]
- Direct buffers / Netty: فریمورکهای شبکه آبجکتهای
DirectByteBufferمیسازند که خارج از heapاند و GC دیر آزادشان میکند؛ با-XX:MaxDirectMemorySizeمحدود کن. - پشتهٔ تردها: ۱۰۰۰ تردِ پلتفرمی × ۱MB = ۱GB حافظهٔ بومی که در
-Xmxاصلاً دیده نمیشود. - glibc malloc arenas: روی لینوکس، تعدادِ زیادِ arena حافظهٔ RSS را باد میکند؛
MALLOC_ARENA_MAX=2یک ترفندِ کلاسیکِ کاهشِ RSS در کانتینر است. - Metaspace: رشدِ بیسقف تا اتمامِ RAM.
برای دیدنِ همهٔ تکهها (نه فقط heap) با -XX:NativeMemoryTracking=summary استارت بزن و بعد jcmd <pid> VM.native_memory summary. این تنها راهِ اثباتِ اینکه «heap سالم است ولی Metaspace/Direct/Thread حافظه را میخورد» است. برای سایزینگِ درستِ کانتینر: -Xmx را روی حدودِ ۷۰–۷۵٪ حافظهٔ کانتینر بگذار (یا -XX:MaxRAMPercentage) و ~۲۵٪ را برای این حافظههای خارج از heap کنار بگذار.
۳) TLAB: تخصیص چطور «فقط جابهجاییِ یک اشارهگر» است
فصل گفت تخصیص در Eden «فقط bump کردنِ یک pointer» است. اما اگر همهٔ تردها روی یک pointerِ مشترک رقابت کنند، باید قفل بگیرند و کند میشود. راهحل: TLAB (Thread-Local Allocation Buffer). هر ترد یک تکهٔ اختصاصی از Eden میگیرد و داخلِ آن بدونِ هیچ همگامسازیای فقط pointer را جلو میبرد. برای همین تخصیصِ آبجکتِ کوچک در جاوا عملاً چند نانوثانیه است.
وقتی آبجکت از TLAB بزرگتر است یا TLAB پر شده، تخصیص به مسیرِ کند (قفلِ مشترک) یا مستقیم به old gen میافتد — «allocation outside TLAB». اگر پروفایلر (async-profiler حالتِ alloc) نرخِ بالای «allocation outside TLAB» نشان داد، یعنی آبجکتهای خیلی بزرگ میسازی. این هم توضیح میدهد چرا escape analysis + scalar replacement انقدر قدرتمند است: آبجکتی که فرار نکند حتی وارد TLAB هم نمیشود.
۴) ارجاعهای بیننسلی: card table و remembered set
اینجا یک تناقضِ ظاهری هست که سنیورها باید حلش را بلد باشند: young GC میخواهد فقط young را بگردد تا سریع باشد؛ اما اگر یک آبجکتِ old به یک آبجکتِ young اشاره کند، آن young زنده است — پس بدونِ اسکنِ کلِ old از کجا بفهمیم؟
راهحل: heap به «کارتهای» ۵۱۲بایتی تقسیم میشود. هر بار که یک فیلدِ ارجاعی را مینویسی (a.f = b)، یک قطعه کدِ ریزِ مخفی به نامِ write barrier آن کارت را «کثیف» علامت میزند. حالا young GC فقط کارتهای کثیف را بهعنوانِ ریشهٔ اضافی میگردد، نه کلِ old را. G1 یک قدم جلوتر میرود: برای هر region یک remembered set (RSet) نگه میدارد که ثبت میکند کدام regionها به این region اشاره دارند — برای همین میتواند یک region را تنها جمع کند.
همان RSetها و write barrierها رایگان نیستند: هم CPU (روی هر نوشتنِ ارجاع) و هم حافظه مصرف میکنند. در برنامههایی با نرخِ خیلی بالای موتیشنِ ارجاع (مثلاً گرافِ آبجکتِ بزرگ و پرتغییر)، سربارِ RSet یکی از دلایلی است که گاهی Parallel GC از G1 throughput بهتری میدهد. مارکِ همزمانِ G1 هم از یک write barrierِ دیگر به نامِ SATB استفاده میکند تا آبجکتی که در حینِ مارک ناپدید میشود را از دست ندهد.
۵) time-to-safepoint: چطور یک ترد کلِ JVM را قفل میکند
فصل گفت GC در safepoint همه را نگه میدارد. نکتهٔ سنیوریِ نادیده: رسیدنِ همه به safepoint فوری نیست. JVM «تعاونی» است؛ هر ترد فقط در نقاطِ خاصی (بازگشتِ متد، لبهٔ برگشتِ حلقه) safepoint را چک میکند. مشکل: HotSpot برای سرعت، در حلقههای شمارشیِ ساده (شمارندهٔ int با کرانِ مشخص) این چک را حذف میکند!
یک حلقهٔ سنگینِ for (int i=0; i<HUGE; i++) بدونِ فراخوانیِ متد ممکن است چند صد میلیثانیه هیچ safepointی نزند. حالا اگر GC یا یک deopt بخواهد شروع شود، باید منتظرِ همهٔ تردها بماند — و این یک ترد کلِ JVM را نگه میدارد: pauseِ GC که باید ۵ms باشد، ناگهان ۳۰۰ms میشود. به این «Time-To-Safepoint spike» میگویند و در لاگ زیرِ [safepoint] (با -Xlog:safepoint) بخشِ reaching بالا میرود، نه بخشِ کارِ واقعیِ GC. علتِ دیگر: page faultِ سنگین یا JNI critical طولانی.
حلقههای شمارشیِ خیلی طولانی را بشکن، یا -XX:+UseCountedLoopSafepoints بده تا در لبهٔ برگشت هم poll بگذارد. اما اول اثبات کن که TTSP مشکل است (-Xlog:safepoint)، نه اینکه کورکورانه فلگ اضافه کنی.
۶) mark word فقط hashCode نیست: حالتهای قفل
فصل mark word را «چیزهای مدیریتی» نامید. بازش کنیم، چون سؤالِ مصاحبه است. همان ۸ بایت، بسته به وضعیت، معنایِ متفاوت دارد و پایینش تگِ حالت است:
- باز (unlocked): hashCode + بیتهای سن (age).
- قفلِ سبک (thin / lightweight): یک CAS اشارهگری به Lock Record روی پشتهٔ ترد میگذارد — قفلِ بیرقابتِ ارزان.
- قفلِ سنگین (inflated / heavyweight): وقتی رقابت پیش میآید، قفل «باد میکند» و mark word به یک
ObjectMonitor(mutex/park سطحِ OS) اشاره میکند. - علامتِ GC: در حینِ جمعآوری.
سالها یک حالتِ چهارم به نامِ biased locking بود (بهینهسازی برای قفلی که همیشه یک ترد میگیرد). در JDK 15 با JEP 374 بهصورتِ پیشفرض غیرفعال و deprecated شد، چون پیچیدگیاش با الگوهای امروزی (thread poolها، lambdaها) نمیارزید و «revokeِ bias» خودش TTSP spike میساخت. سنیورِ بهروز این را میداند: دیگر رویش حساب نکن.
۷) سه نکتهٔ ظریفِ دیگر که فصل باز نکرد
Compressed Class Space — «متاسپیسِ دوم». وقتی compressed class pointers روشن است (heap زیرِ ۳۲GB)، خودِ متادیتای klass در یک ناحیهٔ جدا و پیوسته به نامِ Compressed Class Space ذخیره میشود (پیشفرض ~۱GB رزرو، با -XX:CompressedClassSpaceSize). پس Metaspace در واقع دو تکه است و میتوانی بهطورِ خاص OutOfMemoryError: Compressed class space بگیری (جدا از Metaspace) — معمولاً وقتی هزاران کلاسِ داینامیک تولید میکنی.
در G1 هر آبجکتی که از نصفِ اندازهٔ region بزرگتر باشد «humongous» است، مستقیم در old gen و در regionهای پیوسته تخصیص مییابد. آرایههای بزرگِ اولیه (مثلاً byte[] چند مگابایتی) قاتلِ رایجاند: تخصیص و آزادسازیشان region را تکهتکه میکند و وقتی regionِ پیوستهٔ کافی پیدا نشود، Full GC راه میاندازد. اگر لاگت Pause Full و humongous allocation نشان داد، -XX:G1HeapRegionSize را بزرگتر کن یا آبجکتهای غول را بشکن.
Inline cache و چندریختیِ گران. JIT هر call siteِ مجازی را بر اساسِ تایپهایی که دیده رنگ میکند: monomorphic (یک تایپ → inline + یک guard)، bimorphic (دو تایپ)، و megamorphic (بیش از دو → تسلیم میشود، vtable lookup میکند و inline نمیکند). یک interfaceِ خیلی پرمصرف با دهها پیادهسازی (مثلاً یک hot pathِ لاگینگ یا یک ابسترکشنِ عمومی) میتواند call siteها را megamorphic و کند کند — گاهی «کمتر ابسترکشن» در hot path واقعاً سریعتر است. مرتبط: OSR (On-Stack Replacement) اجازه میدهد یک حلقهٔ طولانی در حالِ اجرا کامپایل شود؛ و intrinsicها متدهایی مثل System.arraycopy، Math.max، Integer.bitCount هستند که JIT با اسمبلیِ دستنویس/دستورِ CPU جایگزین میکند.
۸) startup و warmup: هزینهای که میکروسرویسها را میسوزاند
JIT «گرمشدن» میخواهد: چند ثانیهٔ اول کد تفسیری و کند است. برای یک سرویسِ بلندمدت مهم نیست، اما برای serverless، scale-to-zero و اسکیلِ سریع در Kubernetes فاجعه است. راهکارهای مدرن:
- AppCDS (Application Class Data Sharing): کلاسهای پارسشده را در یک آرشیو ذخیره میکند تا استارتِ بعدی نپارسدشان.
- JEP 483 (AOT Class Loading & Linking، JDK 24، پروژهٔ Leyden): یک قدم جلوتر از AppCDS؛ کلاسها را از قبل load و link هم میکند و در «AOT cache» میگذارد؛ در دموی رسمی استارتِ Spring PetClinic تا ~۴۲٪ سریعتر شد. نیازمندِ یک «training run» شبیهِ پروداکشن است.
- CRaC (Coordinated Restore at Checkpoint): از یک JVMِ گرمشده checkpoint میگیرد و در چند میلیثانیه restore میکند — هم استارتِ فوری، هم peak performanceِ فوری (در بیلدهای Azul/Zulu موجود است).
- GraalVM Native Image: کلِ برنامه را AOT به یک باینریِ بومی کامپایل میکند: استارتِ تقریباً آنی و حافظهٔ کم، اما بدونِ JIT (peakِ کمتر برای بارهای سنگین) و با دنیای بسته (reflection نیاز به کانفیگ دارد).
JIT در طولِ زمان با پروفایلِ واقعی به peak throughputِ بالا میرسد اما گرمشدن میخواهد. AOT (Native Image/Leyden) سریع استارت میزند اما ممکن است سقفِ throughput پایینتری داشته باشد. برای سرویسِ همیشهروشنِ پرترافیک → JIT/HotSpot؛ برای functionِ کوتاهعمر یا scale-to-zero → AOT/CRaC.
فصل گفت compact object headers (JEP 450) در JDK 24 «آزمایشی» است. آپدیت: در JDK 25 با JEP 519 به یک قابلیتِ محصولِ کامل ارتقا یافت (mark word و klass pointer در یک واژهٔ ۶۴بیتیِ واحد ادغام میشوند؛ header از ۱۲ به ۸ بایت، تا ~۲۲٪ صرفهجوییِ heap در بنچمارک). هنوز پیشفرض نیست و با -XX:+UseCompactObjectHeaders روشن میشود، ولی فلگِ experimental حذف شده.
سؤالاتِ سختِ سنیوری (تکمیلی)
نه. volatile فقط دیدهشدن و ترتیب را تضمین میکند (و خواندن/نوشتنِ اتمیکِ long/double)، اما count++ یک read-modify-write است و دو تردِ همزمان میتوانند آپدیتِ هم را گم کنند. برای اتمیکبودنِ خودِ عملیات به CAS نیاز داری: AtomicInteger/AtomicLong (یا زیرِ رقابتِ بالا LongAdder) یا VarHandle. جملهای که باید بگویی: «volatile مشکلِ visibility را حل میکند، نه مشکلِ atomicity را.»
بدونِ volatile، تخصیصِ آبجکت سه مرحله است (تخصیصِ حافظه، اجرای سازنده، انتساب به فیلد) و JMM اجازهٔ جابهجایی اینها را میدهد. تردِ دوم میتواند ارجاعِ غیرnull را ببیند در حالی که سازنده هنوز کامل اجرا نشده — یعنی آبجکتِ نیمهساخته. volatile این reorder را ممنوع و انتشار را امن میکند. (در عمل الگوی Holder را ترجیح بده که این دام را کلاً ندارد.)
kernel بر اساسِ RSSِ کلِ پروسه میکشد، نه heapِ جاوا؛ و RSS = heap + Metaspace + code cache + پشتهٔ تردها + direct buffers + ساختارهای GC + malloc arenas. نبودِ OutOfMemoryError یعنی heap سالم است و مشکل خارج از heap است. با -XX:NativeMemoryTracking=summary + jcmd <pid> VM.native_memory تجزیه کن؛ مظنونها: direct buffers (Netty)، نشتیِ ترد، Metaspace بیسقف، یا malloc arenaها (با MALLOC_ARENA_MAX=2 تست کن). درمان سایزینگ: -Xmx ~۷۵٪ حافظهٔ کانتینر و ~۲۵٪ برای بقیه.
از طریقِ time-to-safepoint. GC باید همهٔ تردها را در safepoint نگه دارد، اما HotSpot در حلقههای شمارشیِ سادهٔ int چکِ safepoint نمیگذارد. یک حلقهٔ داغِ طولانیِ بدونِ فراخوانیِ متد صدها میلیثانیه به safepoint نمیرسد و همه منتظرش میمانند. در -Xlog:safepoint بخشِ «reaching» بالاست نه کارِ GC. درمان: شکستنِ حلقه یا -XX:+UseCountedLoopSafepoints — اما اول اثبات کن.
با card table: heap به کارتهای ۵۱۲بایتی تقسیم میشود و یک write barrier روی هر نوشتنِ فیلدِ ارجاعی کارتِ مربوطه را کثیف میکند. young GC فقط کارتهای کثیفِ old را بهعنوانِ ریشهٔ اضافی میگردد. G1 علاوه بر این برای هر region یک remembered set نگه میدارد. هزینهاش سربارِ write barrier و حافظهٔ RSet است — یکی از دلایلی که Parallel گاهی throughput بهتری از G1 دارد.
mark word حالتِ قفل را کد میکند: unlocked، thin/lightweight (CAS اشارهگر به Lock Record روی پشته، برای قفلِ بیرقابت)، و inflated/heavyweight (اشاره به ObjectMonitor سطحِ OS هنگامِ رقابت). biased locking حالتِ چهارمی بود برای قفلی که همیشه یک ترد میگرفت، اما در JDK 15 (JEP 374) پیشفرض غیرفعال و deprecated شد چون با thread poolها/lambdaهای امروزی نمیصرفید و «bias revocation» خودش safepoint/TTSP میساخت.
JIT هر call siteِ مجازی را بر اساسِ تعدادِ تایپهای دیدهشده رنگ میکند: monomorphic (۱ تایپ → inline + guard)، bimorphic (۲)، megamorphic (بیش از ۲ → inline را رها میکند و vtable lookupِ گران میزند). یک interfaceِ خیلی مشترک با دهها پیادهسازی call siteها را megamorphic میکند و چون inlining از بین میرود، بهینهسازیهای پس از inline هم از بین میروند. در hot path گاهی کاهشِ ابسترکشن (یا سیلکردنِ تایپ) واقعاً سریعتر است.
علت: JIT گرمشدن میخواهد؛ کدِ اولیه تفسیری است. گزینهها: AppCDS (آرشیوِ کلاسهای پارسشده)؛ JEP 483 / AOT cache در JDK 24 که load+link را هم از قبل انجام میدهد (~۴۰٪ استارتِ سریعتر، نیازمندِ training run)؛ CRaC که از JVMِ گرم checkpoint/restore میکند (استارت و peakِ فوری)؛ و GraalVM Native Image که کلِ برنامه را AOT میکند (استارتِ آنی و حافظهٔ کم، اما بدونِ JIT پس peakِ کمتر و reflection نیازمندِ کانفیگ). قاعده: سرویسِ همیشهروشنِ پرترافیک → JIT؛ functionِ کوتاهعمر/scale-to-zero → AOT یا CRaC.
heap تنها تکه از حافظهٔ پروسه است — کانتینر بر اساسِ RSS میکشد، پس NMT بلد باش. درستیِ همزمانی از JMM میآید: volatile = visibility+ordering (نه atomicity)، انتشارِ امن از فیلدِ final/volatile. تخصیص با TLAB سریع است؛ ارجاعِ بیننسلی با card table/RSet؛ و یک حلقهٔ شمارشیِ داغ از راهِ time-to-safepoint کلِ JVM را قفل میکند. biased locking رفته (JDK 15). و برای startup، دنیای مدرن از AppCDS → AOT (Leyden) → CRaC → Native Image انتخاب میکند و آگاهانه JIT-peak را با startup-speed تاخت میزند.
JVM یک کامپیوترِ خیالی است که بایتکدِ .class را اجرا میکند تا کدِ تو همهجا کار کند. ClassLoaderها کلاسها را تنبل و در سه فاز (load → link → init) و با قانونِ «اول از والد بپرس» (برای امنیت) لود میکنند. حافظه در نواحیِ مشخص زندگی میکند: heap برای آبجکتها، stack برای متغیرهای موقتِ هر ترد، metaspace برای متادیتای کلاس. Garbage Collector با پیداکردنِ آبجکتهای غیرقابلدسترس از لنگرها حافظه را خودکار آزاد میکند، و با تکیه بر «بیشترِ آبجکتها جوان میمیرند» heap را نسلی تقسیم میکند. G1 پیشفرضِ همهمنظوره است؛ ZGC برای تأخیرِ فوقکم. JIT کدِ داغ را در لحظه به کدِ ماشین کامپایل میکند (C1 سریع، C2 تهاجمی). و بیشترِ باگهای حافظه در جاوا در واقع نشتی = قابلیتِ دسترسیِ ناخواسته هستند، نه کمبودِ واقعیِ رم.
منابع: JEP 439، JEP 474، JEP 490، Inside.java: Generational ZGC.
This chapter is the heart of Java. Get it right and the rest of the language becomes logical instead of memorized. So don't rush. We go step by step — first grab the idea with a simple analogy, then meet its technical name, then wire it back to real code and interview questions.
First we understand what the JVM is and why it even exists. Then we open its three big parts: (1) how classes get loaded, (2) where memory lives (heap, stack, metaspace), (3) how the execution engine makes code fast (JIT) and keeps memory clean (Garbage Collector). We finish with tuning, memory leaks, reading a GC log, and interview questions.
Part 0 — Three words you must understand from scratch
Before anything, let's unpack three words that repeat throughout the chapter, in plain language. If these click, the rest is easy.
What does "abstraction" mean?
Imagine driving a car. You only deal with the steering wheel, gas pedal, and brake. What explosions happen inside the engine, how fuel burns, how the gears turn — none of that is your concern, and you don't need to know.
That's abstraction: "hide the complexity, expose only what the user needs."
What is a "virtual machine"?
The JVM is an imaginary, made-up computer running inside your real computer (Windows/Mac/Linux). You write your code for this "imaginary computer," not for Windows or Mac. The JVM abstracts (hides) the real hardware's complexity and translates your code into the real hardware's language.
The result: write your code once, run it everywhere — Windows, Mac, Linux, servers, phones. That's Java's famous slogan: Write once, run anywhere.
So when we say the JVM is an "abstract machine," we mean an imaginary computer that hides the underlying hardware details from you.
What does "stack-based" mean?
The JVM uses a structure called a stack to do its computations. A stack is exactly what it sounds like: a pile of plates. The last plate you put on is the first one you take off (this is called LIFO: Last In, First Out).
For example, to compute 2 + 3, the JVM does this: push 2 onto the stack, push 3 on top, then see the add instruction, pop both, add them, and push 5. You don't need to memorize the details right now; just know the stack is the JVM's moment-to-moment workbench.
Abstraction = hiding complexity. Virtual machine = an imaginary computer that runs Java code so you don't depend on the hardware. Stack-based = a moment-to-moment workbench for computation, with "last in, first out" logic.
Part 1 — What exactly does the JVM do? (The big picture)
Let's trace the whole path from "code you write" to "what the CPU does" with an analogy.
You write a recipe (your .java code). This recipe is in human English — readable to you, but hardware doesn't understand it.
A translator (called the compiler, javac) comes and translates your recipe into a standard international language (something like Esperanto). This intermediate language is called bytecode (the .class file).
Why do this? Because every chef understands this intermediate language — it doesn't matter if the chef is French (Windows) or Japanese (Linux). Each has a local translator called the JVM that takes the bytecode and runs it in its own hardware's language.
So the full path is:
your code (.java) ──javac──▶ bytecode (.class) ──JVM──▶ runs on real hardware
readable English standard intermediate language the OS's native machine code
The subtle point: bytecode is platform-independent — the same .class file works on any system. The only thing that differs between operating systems is that last step (the JVM), which has a separate build for each OS.
The JVM does three things with bytecode:
- Load: find the
.classfile and bring it into memory. - Verify: check that the bytecode is safe and well-formed.
- Execute: first interpret it line by line; and for the parts that repeat a lot ("hot code"), compile them to native machine code at runtime so they run fast. This runtime compilation is called JIT (Just-In-Time).
The JVM's three subsystems
- Class Loader = procurement & storage: goes and finds the
.classfiles, brings them in, and checks they're intact. - Runtime Data Areas = the physical kitchen space: the big fridge (heap), the moment-to-moment workbench (stack), the recipe library (metaspace) — where everything lives during work.
- Execution Engine = the head chef and assistants: the one who actually cooks; includes the interpreter, the smart compiler (JIT), and the cleaner (Garbage Collector).
We'll open all three departments below.
The most important mindset shift: "standard" vs "product"
Here's a senior-level point many people get wrong:
The government writes a "building code": "every house must have a door, windows, a roof, and plumbing." The code doesn't say how to build it, only what output it must have. This is Java's Specification.
Now several contractors come and build houses following that code:
- Oracle with a product called HotSpot (the most famous and the default) — packed with clever tricks for speed and memory management.
- IBM with OpenJ9 — focused on lower RAM usage.
- Azul with Zing — focused on very high scalability.
The conclusion here is interview gold:
The JVM Specification = the code/standard (it defines behavior). HotSpot = Oracle's product that implements that standard. Almost everything we say in this chapter about GC, object headers, and JIT describes HotSpot's specific behavior, not a requirement of the spec. A junior thinks "Java = HotSpot." A senior knows Java is a standard and HotSpot is just one implementation of it — so when they hit a problem, they reach for HotSpot's specific tuning knobs rather than assuming it's inherent to the language.
Part 2 — Class loading: loading → linking → initialization
When your program runs, the JVM does not load all classes at once. It loads a class only when you actually need it for the first time. This is called lazy loading.
When you buy a 10-volume encyclopedia, you don't read all 10 volumes right then! You only open the "letter K" volume when you actually need a word starting with K. The JVM is the same: it loads a class exactly when you first reach it (e.g., when you new it or call a static method). This makes the program start faster and avoids wasting memory.
When the JVM finally needs a class, it prepares it in three phases in a precise order (per JLS §12.4 and JVMS §5). Let's see them through the analogy of hiring a new employee.
Phase 1: Loading
The "HR department" (the ClassLoader) reads the bytes of the .class file (from disk, from inside a JAR, from the network, or even bytes generated on the fly) and builds a personnel file in memory — in Java this file is a Class<?> object placed on the heap.
A class's runtime identity equals the combination of (fully-qualified name + the exact ClassLoader that loaded it). Meaning: if the exact same bytes are loaded by two different ClassLoaders, the JVM sees them as two completely different types!
Because of that rule, you might see a strange error: ClassCastException: com.x.Foo cannot be cast to com.x.Foo — "Foo cannot be cast to Foo"! How? Because two different ClassLoaders loaded these two Foos, so to the JVM they're separate types. This trap is common in application servers, OSGi, and hot-reload.
Phase 2: Linking — three sub-steps
Now that the file is built, we connect the employee to the company's systems. Three sub-steps:
Verification — the security check: the bytecode is checked for safety and correctness: is it type-safe? Are the stack operations valid? Does it have illegal jumps into the code? This is the backbone of JVM security — it's this verification that prevents tampered bytecode from corrupting memory.
Preparation — the empty desk:
staticfields are created and given their default value (numbers0, objectsnull, booleansfalse) — not their real value. Like giving the employee an empty desk and a blank notebook, not their actual tools.Resolution — turning promises into addresses: in your code you wrote "go work with the
Databaseclass." For now that's a symbolic name, vague. Here the JVM turns that name into a direct reference (the exact address of that thing in memory). This step can also be lazy and deferred to the last moment.
Phase 3: Initialization
Now it's "the first day of work." The JVM runs a special hidden method called <clinit> (short for class initializer). You don't write this method; the compiler builds it from two things and runs it top to bottom:
- the real initialization of
staticfields (the0s andnulls from the previous step now get their real values), - and the
static { ... }blocks.
When does it run? Only on "active use" of the class: new, calling a static method, reading/writing a static field (unless it's a compile-time constant), reflection, or initializing a subclass.
The JVM guarantees <clinit> runs exactly once and in a fully thread-safe way. So if 100 threads reach this class for the first time simultaneously, the JVM queues them itself: only one initializes, the rest wait until it's ready. Without you writing a single synchronized keyword.
From this guarantee, a very popular and clever pattern for a lazy Singleton is built:
public class Config {
private Config() {} // nobody outside can construct it
// This inner class isn't loaded until getInstance() is called (lazy loading).
private static class Holder {
static final Config INSTANCE = new Config(); // <clinit> once, thread-safe, by the JVM
}
public static Config getInstance() { return Holder.INSTANCE; }
}
Until someone calls getInstance(), the Holder class doesn't even exist (lazy). The moment it's called, the JVM loads Holder, and because it guarantees <clinit> runs exactly once and safely, we've built a Singleton that is lazy, thread-safe, and lock-free — the best of both worlds.
ClassLoader hierarchy and the delegation model
The HR departments (ClassLoaders) work as a parent-child chain with one strict rule: "ask your boss first." This is the parent-delegation model: each loader first asks its parent to load the class; only if the parent fails does it try itself.
Bootstrap ClassLoader (native/C++, loads the Java core: java.lang.* etc.)
└─ Platform ClassLoader (formerly "extension"; the JDK's own standard modules)
└─ System/Application ClassLoader (your program's classpath)
└─ Custom / web-app ClassLoaders (e.g., Tomcat)
Why does this rule exist? Security. Let's see with an example:
Suppose the CEO (Bootstrap) keeps the company's real, official seal in their safe — the world's most trusted seal. A hacker makes a fake seal that looks exactly like it and secretly places it on your desk (your project folder); they name it java.lang.String.
Without the delegation rule: you need a seal, you look at your own desk, grab the fake seal, and stamp the documents with it. The hacker wins — their virus-laced String code runs instead of Java's safe String, and the whole system is compromised.
With the delegation rule: you're not allowed to grab a seal from your desk on your own. You first ask the CEO (Bootstrap). They say, "I have the original String, use this." So the fake seal on your desk is never even seen. The hacker loses.
Delegation prevents anyone from shadowing a core Java class (like java.lang.String) by dropping a same-named file in the project folder. Because the JVM always asks the core and upstream parents first, the original, safe version always wins.
From Java 9 onward, the old "extension" loader became the Platform ClassLoader, and the whole thing is built on the module system (JPMS). Also, because Bootstrap is written in C++ and isn't a Java object, if you call SomeCoreClass.class.getClassLoader() you get null (meaning "I was loaded by Bootstrap").
In one Tomcat, 10 different web apps might run at once, each wanting a different version of a library (e.g., Log4j). If they all asked the parent, they'd collide. So Tomcat deliberately tells each web app's loader: "search your own folder first; if you don't find it, then ask the parent." This is the inverse of standard delegation, and its goal is to isolate apps so they don't break each other.
Part 3 — Runtime data areas: where does memory live?
This part shows you where each thing is stored when your program runs. Let's start with the kitchen analogy, then see the precise table.
- Heap = the big fridge: everything you create with
new(all objects and arrays) is kept here in bulk. This area is managed by the Garbage Collector. - Stack = the head chef's moment-to-moment workbench: each thread's small, temporary things (local variables, references, return address) go here and are cleared when the method finishes.
- Metaspace = the recipe library: information about the classes themselves (not objects) lives here — structure, methods, bytecode.
Now the full table of areas and what error each throws when it fills up:
| Area | Shared among whom? | What it holds | Error on filling |
|---|---|---|---|
| Heap | whole JVM | all objects and arrays | OutOfMemoryError: Java heap space |
| Metaspace | whole JVM | class metadata, method bytecode | OutOfMemoryError: Metaspace |
| JVM Stack | per thread | frames (locals, operands, return address) | StackOverflowError |
| PC Register | per thread | address of the current bytecode instruction | — |
| Native Method Stack | per thread | C/C++ code frames (via JNI) | — |
| Code Cache | whole JVM | JIT-compiled native code | CodeCache is full (JIT turns off) |
Two important notes about this table:
1. Each thread has its own stack. If a method calls itself infinitely (endless recursion), the stack fills up and you get StackOverflowError. On the other hand, if you create tens of thousands of platform threads, the system's native memory runs out (each stack is about 512KB–1MB, tunable with -Xss).
It's exactly this "each thread needs an expensive native stack" pressure that virtual threads solve: they don't permanently pin a native stack, so you can have millions of them. (They get their own chapter.)
2. Metaspace replaced PermGen. This is one of the most important Java 8 changes and a common interview question — we open it fully in the next section.
Zooming in on Metaspace: why and how it replaced PermGen
Picture a car factory:
- Heap = the big parking lot: every car built (say 10,000 of one model) is parked here. In Java, these cars are the objects.
- Metaspace = the blueprint & engineering room: where the car's blueprint is kept.
The key point: to build 10,000 cars you need only one blueprint, not 10,000! In Java it's exactly the same — you make thousands of objects from one class (all in the Heap), but the structural info of the class itself is stored only once in Metaspace.
What exactly is in Metaspace? Metadata (i.e., "data about data"): the class name, the list of methods and their parameters, the list of fields and their types, inheritance info, and the constant pool. Note: the value of a field (e.g., that a user's name is "Ali") is on the Heap; but the fact that "the User class has a name field of type String" is in Metaspace.
Before Java 8 this space was called PermGen, lived inside the Heap, and had a fixed, small cap. If the program loaded many classes, it filled up and you got the infamous OutOfMemoryError: PermGen space. From Java 8, this space moved to the OS's native memory (off-heap) and was renamed Metaspace. The benefit: it now grows dynamically and doesn't crash on a small fixed cap until your physical RAM is full.
Two main causes: (1) ClassLoader leak — especially in web servers: every time you redeploy code, Tomcat creates a new ClassLoader and loads all classes again; if the old ClassLoaders aren't discarded properly, duplicate copies pile up in Metaspace. (2) Dynamic class generation — libraries like Hibernate/Spring/CGLIB generate classes at runtime; if they run away, they fill Metaspace. That's why in production you cap it with -XX:MaxMetaspaceSize.
Part 4 — Object layout and headers (HotSpot, 64-bit)
Now that we know objects live on the Heap, let's see what an object exactly looks like in memory. Besides your own fields, each heap object has a header the JVM needs to manage it:
[ mark word: 8 bytes ][ klass pointer: 4 bytes ][ your fields... ][ padding to a multiple of 8 ]
- Mark word — management stuff: identity hashCode, age bits for GC, lock state, and during a GC move, a forwarding pointer.
- Klass pointer — points to the class metadata in Metaspace (i.e., "which class am I an instance of?").
When the heap is smaller than ~32GB (the default), HotSpot uses compressed oops: it stores pointers in 4 bytes instead of 8 (as a scaled offset). This is a big memory saving — we come back to it in "why 31GB beats 33GB."
So an empty Object is 16 bytes (12 bytes header + 4 bytes padding). Arrays add an extra 4-byte "length" field.
This header overhead explains why boxing is expensive: a boxed Integer takes about 16 bytes (header + value + padding) plus a separate reference, whereas a raw int is just 4 bytes. In hot loops and large data structures, this difference is really felt.
Java 24's JEP 450 (compact object headers, still experimental) brings the header down to 8 bytes. Good to know, but not the default yet.
Part 5 — Garbage Collection fundamentals
Here we reach Java's magic: you don't free memory by hand; an automatic Garbage Collector does it. But how does it know which object is "garbage"?
GC works by "reachability," not counting
Imagine each object is a balloon and references are strings tying the balloons to each other and to a fixed anchor (a GC root). Roots are things that are always "alive": local variables of running threads, static fields, and JNI references.
The GC starts from the anchors and marks any balloon reachable through the strings as "alive." Any balloon with no path to an anchor is garbage and gets freed.
Because it handles cycles correctly. Suppose object A points to B and B points to A, but neither is tied to an anchor — an isolated "island." With naive reference counting, each one's counter is 1, so they're never freed (a leak!). But Java's tracing GC, starting from anchors and never reaching this island, correctly declares both dead.
The generational hypothesis
An empirical observation that the entire design of modern GCs is built on: most objects die young. That is, most objects become useless very soon after being created (like temporary variables inside a method). The GC exploits this fact by splitting the heap:
Young generation: [ Eden | Survivor S0 | Survivor S1 ] Old generation (Tenured)
- New objects are created in Eden (ultra-fast allocation, just bumping a pointer).
- A young GC (minor GC) copies the live objects of Eden and one Survivor into the other Survivor and increments their age. Since it only touches live objects and most are dead, it's very cheap.
- An object that survives enough young cycles (
-XX:MaxTenuringThreshold, up to 15) is promoted/tenured to the Old generation. - A major GC collects the old generation; and a Full GC collects everything (young + old + often metaspace). The Full GC is the expensive one you want to avoid.
So the GC can safely move objects and count anchors, the JVM briefly pauses all application threads at a safe point (safepoint) — this pause is called Stop-the-World (STW). Every GC has some STW. Modern GCs minimize it by doing most of the work concurrently with the program. Latency-sensitive systems live and die by the length and count of these pauses.
Weak references: weak / soft / phantom
Java has a few special reference types that let you tell the GC "this object isn't all that important":
SoftReference— cleared only under memory pressure. Good for caches (but be careful, it can leak).WeakReference— cleared at the next GC once it's only weakly reachable. The basis ofWeakHashMap.PhantomReference— for deterministic cleanup of native resources after an object dies (theCleanerclass, the modern replacement forfinalize()).
Part 6 — The GC landscape: which one, when?
Java has several different Garbage Collectors, and you pick one with a flag. Each is a trade-off between throughput (total work done) and latency (short pauses).
| Collector | Flag | Pause model | Best for | Trade-off |
|---|---|---|---|---|
| Serial | -XX:+UseSerialGC |
fully STW, single-threaded | small heaps, single-CPU container, CLI tools | simplest, lowest overhead; but long pauses at scale |
| Parallel | -XX:+UseParallelGC |
fully STW, multi-threaded | batch jobs, maximizing throughput | best throughput; worst tail latency |
| G1 | -XX:+UseG1GC (default since Java 9) |
mostly-concurrent mark, STW evacuation | general purpose, large heap, latency/throughput balance | you give a pause target (-XX:MaxGCPauseMillis, default 200ms) |
| ZGC | -XX:+UseZGC |
concurrent, sub-1ms pauses | very large heaps (up to terabytes), hard latency SLA | slightly lower throughput, more CPU/memory overhead |
| Shenandoah | -XX:+UseShenandoahGC |
concurrent, sub-1ms pauses | low latency, Red Hat builds | similar to ZGC |
In all modern JDKs — including Java 17 and Java 21 — the default collector is still G1. ZGC and Shenandoah are opt-in. If someone asks "the default in 17 and 21?", the right answer: G1 (which replaced Parallel back in Java 9).
G1 with a bit of detail
G1 (short for Garbage-First) divides the heap into about 2048 equal regions (each 1–32MB). Each region dynamically gets tagged as Eden, Survivor, Old, or Humongous (objects larger than half a region). G1 marks mostly concurrently, then does STW evacuation pauses that pull live objects first out of the regions with the most garbage — hence "Garbage-First." You give a pause target, not an exact layout; G1 itself decides how many regions to collect to hit your target. G1 effectively eliminated the multi-second Full GCs of the old CMS collector (CMS was fully removed in Java 14).
ZGC and Generational ZGC — state the versions precisely
ZGC is a concurrent, region-based, compacting collector that, using tricks called colored pointers and load barriers, moves objects while the program is running. Its headline: pauses are sub-millisecond (typically 0.05–0.5ms) and do not grow with heap size — it scales to multi-terabyte heaps.
- Java 15: ZGC became production-ready (JEP 377).
- Java 21: added Generational ZGC (JEP 439), but opt-in:
-XX:+UseZGC -XX:+ZGenerational. Plain-XX:+UseZGCgave the old non-generational version. - Java 23: generational mode became the default for ZGC (JEP 474), and the
ZGenerationalflag was deprecated. - Java 24: the non-generational mode was fully removed (JEP 490).
Stating these versions precisely shows the interviewer you actually follow the platform.
When the heap is large and you have a hard need for short tail latency (trading systems, ultra-low-latency services) where even G1's ~100–200ms pauses are too much. For general services, stay on G1 — it usually gives better throughput and less memory overhead, and it's the battle-tested default.
Part 7 — JIT compilation, tiered, and inlining
Remember we said the JVM first interprets code and compiles the hot parts? Let's make it precise.
HotSpot first interprets the bytecode and simultaneously profiles it: how many times was each method called? Which branches are taken most? Which types actually show up? Hot methods are turned into native code by two compilers:
- C1 (client): fast compile, light optimization → fast startup.
- C2 (server): slow compile, aggressive optimization (inlining, loop unrolling, dead-code elimination) → peak performance.
The interpreter is like a rookie cook: it reads the bytecode line by line and executes it (normal speed, but it starts right now). The JIT is like a chef who notices which dish has been ordered a hundred times; they pre-convert that recipe into "automatic muscle memory" (native machine code) so next time it cooks at the speed of light.
Tiered compilation (the default) combines both compilers across 5 levels:
Level 0: interpreter
Level 1: C1, no profiling (trivial methods)
Level 2: C1, limited profiling
Level 3: C1, full profiling ← most methods warm up here
Level 4: C2, fully optimized ← the hottest methods reach here
Code starts interpreted, gets compiled by C1 after warming up with profiling, and the truly hot paths "graduate" to C2. Compiled code lives in the code cache.
If the code cache fills, the JIT turns off and everything falls back to the slow interpreter — a sudden performance cliff. The log says CodeCache is full. In very large programs you sometimes need to raise its cap (-XX:ReservedCodeCacheSize).
Inlining — the most valuable optimization
Inlining means replacing a call to a method with the body of that method. Why does it matter? Because once the body is copied in place, all the other optimizations can cross the method boundary and combine. C2 aggressively inlines small, hot methods.
Many people fear adding getters/setters because they think of "method call cost." In practice C2 inlines these tiny methods and they vanish entirely — as if you'd touched the field directly. So don't worry about getter performance.
Escape analysis and scalar replacement
C2 tries to prove whether an object escapes the method/thread that created it (i.e., is seen anywhere outside). If it's proven never to escape:
- Scalar replacement: the object is never allocated on the heap; its fields go straight into registers/the stack. That's why a hot loop creating a temporary
Pointcan produce zero garbage. - Lock elision: if there's a
synchronizedon an object that doesn't escape, that lock is removed entirely.
// C2 can prove p doesn't escape; the temporary object may never reach the heap.
int sumOfSquares(int[] xs, int[] ys) {
int total = 0;
for (int i = 0; i < xs.length; i++) {
Point p = new Point(xs[i], ys[i]); // scalar-replacement candidate
total += p.x * p.x + p.y * p.y;
}
return total;
}
Escape analysis is not a guarantee — it's best-effort and can fail (e.g., when the object is passed to a non-inlined method). Don't design your program's correctness around it. And if you're benchmarking with JMH, know that this is exactly why JMH uses a Blackhole — to stop the compiler from deleting your "seemingly useless" code entirely.
Deoptimization
C2 makes speculative bets (e.g., "this call site has only ever seen ArrayList, so I'll assume it always does and generate optimized code"). If reality one day violates that assumption (suddenly a LinkedList arrives), the JVM throws away that compiled code, temporarily falls back to the interpreter, and may recompile. This shows up as a short performance dip after a new type appears.
Part 8 — Key flags every senior should know
# heap size — in production set min == max to avoid resize pauses and fragmentation
-Xms4g -Xmx4g
# container-aware (on by default since Java 10+): take heap as a percentage of container memory
-XX:MaxRAMPercentage=75.0
# collector choice
-XX:+UseG1GC # default
-XX:+UseZGC # generational from Java 23+
-XX:MaxGCPauseMillis=100
# per-thread stack size
-Xss512k
# Metaspace cap (otherwise it grows until native memory is exhausted)
-XX:MaxMetaspaceSize=256m
# on OOM: take a heap dump and die fast (for later diagnosis)
-XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/dumps
-XX:+ExitOnOutOfMemoryError
# GC log (unified, since Java 9+)
-Xlog:gc*:file=gc.log:time,uptime,level,tags:filecount=5,filesize=20m
If min and max differ, the committed heap keeps shrinking and growing, which itself causes Full GCs and page faults. Fixing the size removes that churn. In containers prefer MaxRAMPercentage so the JVM respects the cgroup limit rather than seeing the whole host's RAM and then getting OOM-killed.
Part 9 — Memory leaks and classifying OutOfMemoryError
In a GC language, a memory leak means unintended reachability: objects you're done with, but which are still reachable from a GC root, so the GC isn't allowed to free them and they pile up.
Classic leak sources:
staticcollections that only grow (static Map cache = ...that never removes anything).- Unbounded caches — use a size cap instead (Caffeine, LRU).
- Listeners/callbacks that are never unregistered — the subject keeps the observer forever.
ThreadLocalin thread pools — a pooled thread never dies, so theThreadLocalvalue is never freed; alwaysremove()it in afinally.- ClassLoader leak — one lingering reference to a web app's classes keeps the whole ClassLoader (and all its classes in Metaspace) alive across redeploys.
Types of OutOfMemoryError (each means something different)
| Message | Meaning | Common cause |
|---|---|---|
Java heap space |
heap full, GC can't free enough | leak or too-small heap |
GC overhead limit exceeded |
>98% of time in GC, <2% freed | heap nearly full, thrashing |
Metaspace |
class metadata space exhausted | classloader leak, too many dynamic classes |
unable to create new native thread |
OS native memory or ulimit full | thread leak, big -Xss × many threads |
Direct buffer memory |
off-heap direct ByteBuffer exhausted |
leaking direct buffers |
Requested array size exceeds VM limit |
requested an array near Integer.MAX_VALUE |
logic bug |
It's an Error, not an Exception. Catching it is almost always wrong, because the JVM may be in an unrecoverable state. Let the program die fast and find the bug with a heap dump.
Part 10 — How to read a GC log
Modern JVMs (9+) use unified logging (-Xlog:gc*). A G1 young pause looks like this:
[2.335s][info][gc,start] GC(12) Pause Young (Normal) (G1 Evacuation Pause)
[2.340s][info][gc ] GC(12) Pause Young (Normal) (G1 Evacuation Pause) 512M->48M(1024M) 5.219ms
Read it like this: GC number 12 was a young evacuation pause; the heap went from 512M before → 48M after (so 464M was freed), the total heap is 1024M, and the pause took 5.219ms.
Look for these signs:
- Rising count and length of pauses = trouble ahead.
- Seeing
Pause Fullon G1/ZGC = a red flag; investigate (allocation spike, humongous objects, too-small heap). - The live size after each Full GC continually rising = a leak (each collection frees less; the floor keeps rising).
to-space exhausted= young objects are promoting too fast; tune the sizes.
To inspect the heap's contents, take a dump (-XX:+HeapDumpOnOutOfMemoryError or jmap -dump) and open it in Eclipse MAT — its "dominator tree" and "leak suspects" tell you exactly what's holding memory. For live diagnosis: jstat -gcutil <pid> 1s, jcmd <pid> GC.heap_info, and JFR (-XX:StartFlightRecording).
Part 11 — Common pitfalls and subtleties
- Calling
System.gc()— it's a request, not a command; it usually triggers an expensive Full GC and hurts. Neutralize it with-XX:+DisableExplicitGC. finalize()— deprecated, unpredictable, can resurrect an object and delay collection. UseCleaneror try-with-resources.- Object pooling for small objects — allocation in the young generation is nearly free; pooling usually increases GC pressure by keeping objects alive into the old generation. Only pool genuinely expensive resources (connections, threads).
- A huge
-Xmx— a bigger heap means longer pauses and later-but-worse Full GCs. Bigger isn't always better.
Above ~32GB, the JVM can no longer use compressed oops, so every reference doubles from 4 to 8 bytes. This overhead is so large that a 31GB heap can hold more usable objects than a 33GB heap! So either stay under 32GB, or if you cross it, cross it by enough to be worth it.
Part 12 — Interview questions (with answers)
G1 in both. It's been the default since Java 9 (replaced Parallel). ZGC and Shenandoah are opt-in.
It was single-generation until Java 20. Generational ZGC arrived in Java 21 (JEP 439) but was opt-in with -XX:+UseZGC -XX:+ZGenerational. It became the default for ZGC in Java 23 (JEP 474), and non-generational mode was removed in Java 24 (JEP 490).
PermGen (before Java 8) kept class metadata in a fixed-size region of the heap and caused frequent OutOfMemoryError: PermGen space on redeploys. Java 8 replaced it with Metaspace in native memory that grows dynamically. You can still leak it (classloader leak); cap it with -XX:MaxMetaspaceSize.
Most objects die young. So the heap is split, and the young GC uses a copying collector that only touches live objects (copies survivors and resets Eden in one shot). Since most young objects are dead, it does little work and compacts for free — no fragmentation, cost proportional to survivors not garbage.
A safepoint is a point where a thread's state is consistent and known to the JVM so it can be safely paused. Even ZGC needs short STW pauses (e.g., start/end of root marking) to build a consistent snapshot; but the bulk of marking/moving is concurrent. "Concurrent" means most of the work overlaps with the app, not zero pauses.
Via colored pointers (metadata bits inside the pointer itself) plus load barriers: when the app loads a reference, the barrier checks/fixes it, letting ZGC move objects concurrently with the app. The pause work is proportional to the number of GC roots (bounded), not the heap size — so as the heap grows to terabytes, pauses stay flat.
C2 analyzes whether an object escapes the method/thread. If not: scalar replacement (the object isn't allocated on the heap; fields become registers/stack → zero garbage) and lock elision (removing useless synchronized). Best-effort, not a guarantee.
The interpreter runs first while profiling. C1 (fast, light optimization) compiles warm methods with full profiling (level 3); the hottest graduate to C2 (slow, aggressive optimization → level 4). This balances fast startup (C1) with peak throughput (C2).
class Cache {
private static final Map<Key, Value> M = new HashMap<>();
static void put(Key k, Value v) { M.put(k, v); } // never evicts
}
A static Map that only grows is forever reachable from a GC root → unbounded heap growth → OutOfMemoryError: Java heap space. Fix: bound it (Caffeine/LRU), or use WeakHashMap/SoftReference if entries should be collectible.
Integer a = 127, b = 127;
Integer c = 128, d = 128;
System.out.println((a == b) + " " + (c == d));
Prints: true false. Integer.valueOf caches the range −128..127, so a and b are the same cached object (== true), but 128 is outside the cache → two distinct objects (== false). This is autoboxing + the Integer cache, and precisely why you always use .equals() for boxed types.
A class's runtime identity equals (name, defining ClassLoader). If two ClassLoaders each load com.x.Foo, the JVM sees them as two different types, and casting one to the other throws ClassCastException, even if the source is identical. Common in app servers, OSGi, and hot-reload.
The classic signature of a memory leak: each Full GC frees less and the live floor rises. Take a heap dump, open it in Eclipse MAT, and use the dominator tree / leak suspects to find the retaining root — often a static collection, an unbounded cache, or a ThreadLocal not remove()d in a pool.
Above ~32GB, the JVM disables compressed oops, so every reference grows from 4 to 8 bytes. This overhead can reduce usable capacity so much that a 31GB heap holds more live objects than a 33GB one. Only cross the boundary when you truly need to.
No — it's a hint the JVM can ignore, and -XX:+DisableExplicitGC makes it a no-op. When honored it usually forces an expensive Full STW GC, so relying on it in application code is an anti-pattern.
Local primitive variables and object references live on the thread's stack (in the frame); the objects themselves on the heap. Class metadata and static fields are in Metaspace (native), while the object a static field points to is still on the heap. The crisp summary: "primitives and references on the stack, objects on the heap" — with the caveat that escape analysis can keep some short-lived objects off the heap entirely.
Senior notes & advanced edge cases
So far we know where memory lives and how GC cleans up. But in production, the thing that pages you at 3 a.m. is usually the hidden layer underneath those ideas: why your container gets OOM-killed even though -Xmx was respected, why one innocent thread freezes the whole JVM, why "correct" code sees stale values across cores. This section is exactly that layer.
(1) The Java Memory Model — happens-before, volatile, safe publication; (2) RSS vs heap and why containers get killed; (3) TLAB and the real allocation path; (4) card tables & remembered sets for cross-generational references; (5) time-to-safepoint and the counted-loop trap; (6) lock states in the mark word; (7) compressed class space, humongous objects, inline caches; (8) startup & warmup (AppCDS/AOT/CRaC/GraalVM); and 8 hard senior interview questions at the end.
1) The Java Memory Model (JMM): the piece the main chapter skipped
The chapter was about "memory" — but where's its most critical piece, the memory model? The heap is shared across threads, yet each core has its own caches and write buffers, and the compiler/CPU are allowed to reorder instructions. So this question is not obvious at all: "if thread A writes a field, when does thread B see it?"
A writes in his private notebook (the core's cache) and assumes B saw it. But until he copies it onto the shared whiteboard (main memory), B keeps reading the stale version. The JMM is the set of rules for when you must copy to the whiteboard.
The heart of the JMM is the happens-before relation: if action X happens-before action Y, then X's effects are guaranteed visible to Y. The edges that matter:
- Monitor lock: an
unlockon a lock happens-before every subsequentlockon the same lock. - volatile: a write to a
volatilefield happens-before every subsequent read of that field. - Thread start/join:
thread.start()happens-before code inside that thread; that thread's code happens-before a returningthread.join(). - Transitivity: if A→B and B→C then A→C.
Gives: visibility, ordering (no reordering across it), and atomic reads/writes of long/double. Denies: atomicity of compound operations. count++ on a volatile field is still racy, because it's a read-modify-write. For counters use AtomicInteger/LongAdder or a VarHandle (CAS).
Unsynchronized concurrent access where at least one is a write = a data race, and its behavior is undefined — you might not just read a stale value, you might never see the update at all (the compiler can hoist the field into a register and your loop spins forever). This bug "works" on your x86 laptop and explodes on an ARM server or under heavy load — the worst kind of bug.
Safe publication and final fields. How do you hand an object to another thread without a lock? The JMM guarantees that if an object is properly constructed (i.e., this does not escape during the constructor), its final fields are visible to all threads immediately after the constructor returns — this "final-field freeze" is exactly what makes String and immutable objects safe: you can pass them between threads with zero synchronization. Other safe ways: write through a volatile/AtomicReference field, use a static initializer (the <clinit> guarantee), or pass through a concurrent collection.
// Correct double-checked locking: instance MUST be volatile,
// otherwise another thread can see a non-null but half-constructed reference.
class Lazy {
private static volatile Lazy instance; // broken without volatile
static Lazy get() {
Lazy r = instance;
if (r == null) synchronized (Lazy.class) {
r = instance;
if (r == null) instance = r = new Lazy();
}
return r;
}
}
In practice, rarely hand-write DCL — the Holder pattern (from the chapter) is simpler and unbreakable. Know DCL so you can catch the forgotten-volatile trap in code review.
2) RSS vs heap: why a -Xmx2g container gets killed at 3 GB
The most common "production mystery": live heap is 1.2 GB, you set -Xmx2g, yet the container hits OOMKilled (exit 137) at 3 GB. Why? Because the heap is only one slice of the whole process's memory. The kernel kills based on RSS (the process's total resident physical memory), not the Java heap.
A one-line decomposition; RSS is the sum of all of these, not just the heap:
flowchart TD
RSS["Container RSS (what the kernel OOM-kills on)"] --> Heap["Java Heap (-Xmx)"]
RSS --> Meta["Metaspace + Compressed Class Space"]
RSS --> Code["JIT Code Cache"]
RSS --> Stacks["Thread stacks (N x -Xss)"]
RSS --> Direct["Direct / mapped ByteBuffers (NIO, Netty)"]
RSS --> GC["GC structures (card table, RSets, mark bitmaps)"]
RSS --> Native["Native libs, JNI, malloc arenas"]
- Direct buffers / Netty: networking frameworks allocate
DirectByteBuffers that live off-heap and are freed lazily by GC; cap them with-XX:MaxDirectMemorySize. - Thread stacks: 1000 platform threads x 1 MB = 1 GB of native memory that
-Xmxnever counts. - glibc malloc arenas: on Linux a large number of arenas bloats RSS;
MALLOC_ARENA_MAX=2is the classic container RSS-reduction trick. - Metaspace: uncapped, grows until RAM is gone.
To see all the slices (not just heap), start with -XX:NativeMemoryTracking=summary, then jcmd <pid> VM.native_memory summary. This is the only way to prove "the heap is healthy but Metaspace/Direct/Thread memory is eating RAM." For correct container sizing: set -Xmx to ~70–75% of container memory (or -XX:MaxRAMPercentage) and leave ~25% for this off-heap memory.
3) TLAB: how allocation is "just bumping a pointer"
The chapter said Eden allocation is "just bumping a pointer." But if every thread contended on one shared pointer they'd need a lock and it would be slow. The fix: the TLAB (Thread-Local Allocation Buffer). Each thread gets its own private chunk of Eden and inside it just bumps a pointer with zero synchronization. That's why allocating a small object in Java is effectively a few nanoseconds.
When an object is bigger than the TLAB or the TLAB is full, allocation falls to the slow path (shared lock) or straight into the old gen — an "allocation outside TLAB." If a profiler (async-profiler in alloc mode) shows a high "allocation outside TLAB" rate, you're creating very large objects. This also explains why escape analysis + scalar replacement is so powerful: an object that never escapes never even touches a TLAB.
4) Cross-generational references: card tables and remembered sets
Here's an apparent contradiction seniors must be able to resolve: a young GC wants to scan only young to stay fast; but if an old object points to a young object, that young object is alive — so how do we know without scanning the whole old gen?
The fix: the heap is divided into 512-byte "cards." Every time you write a reference field (a.f = b), a tiny hidden snippet called the write barrier marks that card "dirty." Now the young GC only scans the dirty cards as extra roots, not the entire old gen. G1 goes one step further: it keeps a per-region remembered set (RSet) recording which regions point into this region — which is what lets it collect a single region in isolation.
Those RSets and write barriers aren't free: they cost both CPU (on every reference write) and memory. In apps with a very high reference-mutation rate (a large, churny object graph), RSet overhead is one reason Parallel GC sometimes beats G1 on throughput. G1's concurrent marking also uses a different write barrier called SATB (snapshot-at-the-beginning) so it doesn't lose an object that disappears mid-mark.
5) Time-to-safepoint: how one thread freezes the whole JVM
The chapter said GC pauses everyone at a safepoint. The overlooked senior nuance: reaching a safepoint is not instant. The JVM is "cooperative"; each thread only checks for a safepoint at specific points (method returns, loop back-edges). The catch: for speed, HotSpot omits that check in simple counted loops (an int counter with a known bound)!
A heavy for (int i=0; i<HUGE; i++) with no method call inside may go hundreds of milliseconds without hitting a single safepoint. Now if a GC or a deopt wants to start, it must wait for all threads — and this one thread stalls the whole JVM: a GC pause that should be 5 ms suddenly becomes 300 ms. This is a "time-to-safepoint spike," and in the log (-Xlog:safepoint) it shows up in the reaching portion, not in the actual GC work. Other causes: heavy page faults or a long JNI critical section.
Break up very long counted loops, or pass -XX:+UseCountedLoopSafepoints so a poll is placed on the back-edge too. But first prove TTSP is the problem (-Xlog:safepoint) rather than blindly adding flags.
6) The mark word isn't just hashCode: lock states
The chapter called the mark word "management stuff." Let's unpack it, because it's an interview question. Those same 8 bytes mean different things depending on state, with a state tag at the bottom:
- Unlocked: hashCode + age bits.
- Thin / lightweight lock: a CAS installs a pointer to a Lock Record on the thread's stack — a cheap uncontended lock.
- Inflated / heavyweight lock: under contention the lock "inflates" and the mark word points to an OS-level
ObjectMonitor(mutex/park). - GC-marked: during collection.
For years there was a fourth state, biased locking (an optimization for a lock always taken by one thread). In JDK 15, JEP 374 disabled it by default and deprecated it, because its complexity no longer paid off with modern patterns (thread pools, lambdas) and "bias revocation" itself caused TTSP spikes. An up-to-date senior knows: don't rely on it anymore.
7) Three more subtleties the chapter didn't open
Compressed Class Space — the "second Metaspace." When compressed class pointers are on (heap under 32 GB), the klass metadata itself is stored in a separate, contiguous region called the Compressed Class Space (default ~1 GB reserved, tuned with -XX:CompressedClassSpaceSize). So Metaspace is actually two parts, and you can specifically get OutOfMemoryError: Compressed class space (distinct from Metaspace) — usually when you generate thousands of dynamic classes.
In G1 any object larger than half a region is "humongous," allocated directly in the old gen across contiguous regions. Large primitive arrays (a multi-MB byte[]) are the usual culprits: allocating and freeing them fragments regions, and when enough contiguous regions can't be found, it triggers a Full GC. If your log shows Pause Full alongside humongous allocations, increase -XX:G1HeapRegionSize or break up the giant objects.
Inline caches and expensive polymorphism. The JIT colors each virtual call site by the types it has seen: monomorphic (one type → inline + a guard), bimorphic (two types), and megamorphic (more than two → it gives up, does a vtable lookup, and does not inline). A very heavily shared interface with dozens of implementations (a hot-path logging call, a generic abstraction) can turn call sites megamorphic and slow — sometimes "less abstraction" in a hot path really is faster. Related: OSR (On-Stack Replacement) lets a long-running loop be compiled while it's still executing; and intrinsics are methods like System.arraycopy, Math.max, Integer.bitCount that the JIT replaces with hand-written assembly/CPU instructions.
8) Startup & warmup: the cost that burns microservices
The JIT needs to "warm up": the first few seconds run interpreted and slow. For a long-lived service that's irrelevant, but for serverless, scale-to-zero, and fast Kubernetes scaling it hurts. The modern options:
- AppCDS (Application Class Data Sharing): archives parsed classes so the next start doesn't re-parse them.
- JEP 483 (AOT Class Loading & Linking, JDK 24, Project Leyden): one step past AppCDS — it also loads and links classes ahead of time into an "AOT cache"; the official demo shows Spring PetClinic starting up to ~42% faster. Requires a "training run" that mimics production.
- CRaC (Coordinated Restore at Checkpoint): checkpoints an already-warmed JVM and restores it in milliseconds — instant startup and instant peak performance (available in Azul/Zulu builds).
- GraalVM Native Image: compiles the whole program AOT into a native binary: near-instant startup and low memory, but no JIT (lower peak for heavy loads) and a closed world (reflection needs configuration).
The JIT reaches high peak throughput over time using a real profile, but needs warmup. AOT (Native Image/Leyden) starts fast but may have a lower throughput ceiling. For an always-on, high-traffic service → JIT/HotSpot; for a short-lived function or scale-to-zero → AOT/CRaC.
The chapter said compact object headers (JEP 450) were "experimental" in JDK 24. Update: in JDK 25, JEP 519 promoted them to a full product feature (the mark word and klass pointer are merged into a single 64-bit word; the header drops from 12 to 8 bytes, up to ~22% heap savings in benchmarks). It's still not the default and is turned on with -XX:+UseCompactObjectHeaders, but the experimental flag is gone.
Hard senior interview questions (additional)
No. volatile guarantees only visibility and ordering (plus atomic reads/writes of long/double), but count++ is a read-modify-write, so two concurrent threads can lose each other's update. For the operation itself to be atomic you need CAS: AtomicInteger/AtomicLong (or LongAdder under high contention) or a VarHandle. The line to say: "volatile solves visibility, not atomicity."
Without volatile, object construction is three steps (allocate memory, run the constructor, assign to the field) and the JMM permits reordering them. A second thread can see the non-null reference while the constructor hasn't finished — a half-constructed object. volatile forbids that reorder and makes the publication safe. (In practice prefer the Holder pattern, which doesn't have this trap at all.)
The kernel kills based on the whole process's RSS, not the Java heap; and RSS = heap + Metaspace + code cache + thread stacks + direct buffers + GC structures + malloc arenas. No OutOfMemoryError means the heap is healthy and the problem is off-heap. Break it down with -XX:NativeMemoryTracking=summary + jcmd <pid> VM.native_memory; suspects: direct buffers (Netty), a thread leak, uncapped Metaspace, or malloc arenas (test with MALLOC_ARENA_MAX=2). Sizing fix: -Xmx ~75% of container memory, ~25% for the rest.
Via time-to-safepoint. GC must pause all threads at a safepoint, but HotSpot places no safepoint check inside simple int counted loops. A long, hot loop with no method call goes hundreds of milliseconds without reaching a safepoint, and everyone waits for it. In -Xlog:safepoint the "reaching" portion is high, not the GC work. Cure: break up the loop or -XX:+UseCountedLoopSafepoints — but prove it first.
Via the card table: the heap is split into 512-byte cards and a write barrier dirties the relevant card on every reference-field write. The young GC scans only the dirty old-gen cards as extra roots. G1 additionally keeps a per-region remembered set. The cost is write-barrier overhead and RSet memory — one reason Parallel sometimes out-throughputs G1.
The mark word encodes the lock state: unlocked, thin/lightweight (CAS a pointer to a stack Lock Record, for uncontended locks), and inflated/heavyweight (points to an OS-level ObjectMonitor under contention). Biased locking was a fourth state for a lock always taken by one thread, but in JDK 15 (JEP 374) it was disabled by default and deprecated because it no longer paid off with modern thread pools/lambdas and "bias revocation" itself caused safepoint/TTSP spikes.
The JIT colors each virtual call site by how many types it has seen: monomorphic (1 type → inline + guard), bimorphic (2), megamorphic (more than 2 → it abandons inlining and does an expensive vtable lookup). A very heavily shared interface with dozens of implementations makes call sites megamorphic, and because inlining is lost, all the post-inline optimizations are lost too. In a hot path, reducing abstraction (or sealing the types) can genuinely be faster.
The cause: the JIT needs to warm up; early code runs interpreted. Options: AppCDS (archive parsed classes); JEP 483 / AOT cache in JDK 24, which also load+links ahead of time (~40% faster startup, needs a training run); CRaC, which checkpoints/restores a warmed JVM (instant startup and peak); and GraalVM Native Image, which AOT-compiles the whole app (instant startup, low memory, but no JIT so lower peak and reflection needs config). Rule of thumb: always-on high-traffic service → JIT; short-lived/scale-to-zero function → AOT or CRaC.
The heap is only one slice of process memory — the container kills on RSS, so learn NMT. Concurrency correctness comes from the JMM: volatile = visibility+ordering (not atomicity), safe publication via final/volatile fields. Allocation is fast thanks to TLABs; cross-generational references are tracked by card tables/RSets; and a single hot counted loop can freeze the whole JVM via time-to-safepoint. Biased locking is gone (JDK 15). And for startup, the modern world chooses along AppCDS → AOT (Leyden) → CRaC → Native Image, deliberately trading JIT peak for startup speed.
The JVM is an imaginary computer that runs .class bytecode so your code works everywhere. ClassLoaders load classes lazily, in three phases (load → link → init), following the "ask your parent first" rule (for security). Memory lives in defined areas: heap for objects, stack for each thread's temporary variables, metaspace for class metadata. The Garbage Collector frees memory automatically by finding objects unreachable from the roots, and splits the heap generationally by relying on "most objects die young." G1 is the general-purpose default; ZGC is for ultra-low latency. The JIT compiles hot code to machine code on the fly (C1 fast, C2 aggressive). And most memory bugs in Java are really leaks = unintended reachability, not a genuine shortage of RAM.
Sources: JEP 439, JEP 474, JEP 490, Inside.java: Generational ZGC.