Java Core · جاوا پایه سنیورSenior ~65 دقیقه مطالعه~56 min read

درونیات JVM، حافظه و Garbage CollectionJVM Internals, Memory & Garbage Collection

از صفر: JVM چیست، حافظه کجا زندگی می‌کند، GC چطور تمیز می‌کند و JIT چطور سریع می‌کند — با تشبیه‌های واقعی، کدِ اجراشدنی و سؤالات مصاحبه.From zero: what the JVM is, where memory lives, how GC cleans up and JIT speeds things up — with real-world analogies, runnable code, and interview questions.


این فصل، قلب جاواست. اگر این را خوب بفهمی، بقیهٔ زبان برایت «منطقی» می‌شود نه «حفظی». پس عجله نکن؛ قدم‌به‌قدم می‌رویم — اول با یک تشبیه ساده مفهوم را می‌گیریم، بعد واژهٔ فنی‌اش را می‌بینیم، و آخر سرش وصل می‌کنیم به کد واقعی و سؤال مصاحبه.

نقشهٔ راه این فصل

اول می‌فهمیم JVM چیست و چرا اصلاً وجود دارد. بعد سه بخش بزرگش را باز می‌کنیم: (۱) چطور کلاس‌ها بارگذاری می‌شوند، (۲) حافظه کجا زندگی می‌کند (heap، stack، metaspace)، (۳) موتور اجرا چطور کد را سریع می‌کند (JIT) و حافظه را تمیز می‌کند (Garbage Collector). در آخر می‌رسیم به تیونینگ، نشتی حافظه، خواندن لاگ GC و سؤالات مصاحبه.


بخش ۰ — سه واژه‌ای که باید از صفر بفهمی

قبل از هر چیز، سه کلمه‌ای که در کل فصل تکرار می‌شوند را با زبان آدمیزاد باز کنیم. اگر این‌ها جا بیفتند، بقیه راحت است.

«انتزاع» (Abstraction) یعنی چه؟

فرض کن رانندگی می‌کنی. تو فقط با فرمان، پدال گاز و ترمز کار داری. اینکه داخل موتور چه انفجارهایی می‌افتد، بنزین چطور می‌سوزد و چرخ‌دنده‌ها چطور می‌چرخند، اصلاً به تو ربطی ندارد و لازم نیست بدانی.

این یعنی انتزاع: «پیچیدگی را قایم کن، فقط چیزی را نشان بده که کاربر لازم دارد.»

«ماشین مجازی» (Virtual Machine) یعنی چه؟

یک کامپیوترِ خیالی داخل کامپیوتر واقعی

‏JVM یک کامپیوترِ فرضی و ساختگی است که داخل کامپیوتر واقعی تو (ویندوز/مک/لینوکس) اجرا می‌شود. تو کدت را برای این «کامپیوتر خیالی» می‌نویسی، نه برای ویندوز یا مک. JVM پیچیدگی‌های سخت‌افزار واقعی را انتزاع (قایم) می‌کند و کد تو را به زبان سخت‌افزار واقعی ترجمه می‌کند.

نتیجه: یک بار کد بنویس، همه‌جا اجرا کن — روی ویندوز، مک، لینوکس، سرور، موبایل. این همان شعار معروف جاوا است: Write once, run anywhere.

پس وقتی می‌گوییم JVM یک «ماشینِ انتزاعی» است، یعنی یک کامپیوتر خیالی که جزئیات سخت‌افزار زیرش را از تو پنهان کرده.

«پشته‌محور» (Stack-based) یعنی چه؟

‏JVM برای انجام محاسبات از یک ساختار به اسم پشته (stack) استفاده می‌کند. پشته یعنی دقیقاً همان چیزی که از اسمش پیداست: یک دستهٔ بشقاب. آخرین بشقابی که می‌گذاری، اولین بشقابی است که برمی‌داری (به این می‌گویند LIFO: آخرین‌ورودی، اولین‌خروجی).

مثلاً برای محاسبهٔ 2 + 3، JVM این‌طور عمل می‌کند: 2 را روی پشته می‌گذارد، 3 را روی‌اش می‌گذارد، بعد دستور add را می‌بیند، هر دو را برمی‌دارد، جمع می‌کند و 5 را روی پشته می‌گذارد. لازم نیست الان جزئیاتش را حفظ کنی؛ فقط بدان پشته میز کارِ لحظه‌ایِ JVM است.

سه واژهٔ کلیدی

انتزاع = قایم‌کردن پیچیدگی. ماشین مجازی = کامپیوتر خیالی که کد جاوا را اجرا می‌کند تا به سخت‌افزار وابسته نباشی. پشته‌محور = میز کارِ لحظه‌ای برای محاسبات، با منطق «آخرین‌ورودی، اولین‌خروجی».


بخش ۱ — JVM دقیقاً چه می‌کند؟ (تصویر بزرگ)

بیا کل مسیر از «کدی که تو می‌نویسی» تا «کاری که پردازنده انجام می‌دهد» را با یک تشبیه ببینیم.

رستوران بین‌المللی

تو یک دستور پخت می‌نویسی (کد .java تو). این دستور به زبان انگلیسیِ آدم‌ها است — خوانا برای تو، ولی سخت‌افزار آن را نمی‌فهمد.

یک مترجم (به اسم کامپایلر، همان javac) می‌آید و دستور تو را به یک زبان بین‌المللیِ استاندارد ترجمه می‌کند (چیزی شبیه اسپرانتو). به این زبانِ میانی می‌گویند بایت‌کد (فایل .class).

چرا این کار را می‌کند؟ چون این زبان میانی را همهٔ آشپزها می‌فهمند — فرقی نمی‌کند آشپز فرانسوی باشد (ویندوز) یا ژاپنی (لینوکس). هر کدام یک مترجمِ محلی به اسم JVM دارند که بایت‌کد را می‌گیرد و به زبان سخت‌افزارِ خودش اجرا می‌کند.

پس مسیر کامل این است:

کد تو (.java)  ──javac──▶  بایت‌کد (.class)  ──JVM──▶  اجرا روی سخت‌افزار واقعی
   انگلیسیِ خوانا           زبان میانی استاندارد        زبان ماشینِ همان سیستم‌عامل

نکتهٔ ظریف: بایت‌کد مستقل از پلتفرم است — یعنی همان فایل .class روی هر سیستمی کار می‌کند. کاری که سیستم‌عامل‌ها را متفاوت می‌کند، فقط آن آخرین قدم (JVM) است که برای هر سیستم‌عامل نسخهٔ جداگانه دارد.

JVM با بایت‌کد سه کار می‌کند:

  1. بارگذاری (load): فایل .class را پیدا و وارد حافظه می‌کند.
  2. وارسی (verify): چک می‌کند بایت‌کد سالم و امن باشد.
  3. اجرا: ابتدا آن را خط‌به‌خط تفسیر (interpret) می‌کند؛ و بخش‌هایی که خیلی تکرار می‌شوند («کد داغ») را در لحظهٔ اجرا به کد ماشینِ بومی کامپایل می‌کند تا سریع شوند. به این کامپایلِ حین‌اجرا می‌گویند JIT (Just-In-Time).

سه زیرسیستم اصلی JVM

سه دپارتمانِ یک رستوران بزرگ
  • بارگذار کلاس (Class Loader) = بخش تدارکات و انبار: می‌رود فایل‌های .class را پیدا می‌کند، می‌آورد داخل، و چک می‌کند سالم باشند.
  • نواحی دادهٔ زمان‌اجرا (Runtime Data Areas) = فضای فیزیکی آشپزخانه: یخچال بزرگ (heap)، میز کارِ لحظه‌ای (stack)، کتابخانهٔ دستورپخت‌ها (metaspace) — جایی که همه‌چیز حین کار زندگی می‌کند.
  • موتور اجرا (Execution Engine) = سرآشپز و دستیارانش: کسی که واقعاً غذا را می‌پزد؛ شامل مفسر (interpreter)، کامپایلر هوشمند (JIT) و نظافتچی (Garbage Collector).

در ادامه هر سه دپارتمان را دقیق باز می‌کنیم.

مهم‌ترین تغییرِ نگاه: «استاندارد» در برابر «محصول»

اینجا یک نکتهٔ سنیوری است که خیلی‌ها اشتباه می‌فهمند:

قانون ساختمان‌سازی در برابر شرکت پیمانکار

دولت یک «قانون ساختمان‌سازی» می‌نویسد: «هر خانه باید در، پنجره، سقف و لوله‌کشی داشته باشد.» این قانون نمی‌گوید چطور بسازی، فقط می‌گوید چه خروجی‌ای باید داشته باشد. این همان مشخصات (Specification) جاواست.

حالا چند شرکت پیمانکار می‌آیند و طبق این قانون خانه می‌سازند:

  • Oracle با محصولی به اسم HotSpot (معروف‌ترین و پیش‌فرض) — پر از تکنیک‌های خفن برای سرعت و مدیریت حافظه.
  • IBM با محصول OpenJ9 — تمرکز روی مصرف رم کمتر.
  • Azul با محصول Zing — تمرکز روی مقیاس‌پذیریِ بسیار بالا.

نتیجه‌گیری این بخش، که برای مصاحبه طلاست:

استاندارد ≠ پیاده‌سازی

JVM Specification = قانون و استاندارد (رفتار را تعریف می‌کند). HotSpot = محصولِ Oracle که آن قانون را پیاده کرده. تقریباً هر چیزی که در این فصل دربارهٔ GC، هدرِ آبجکت و JIT می‌گوییم، مربوط به رفتار خاصِ HotSpot است، نه الزامِ استاندارد. یک جونیور فکر می‌کند «جاوا = HotSpot». یک سنیور می‌داند جاوا یک استاندارد است و HotSpot فقط یکی از پیاده‌سازی‌های آن — پس وقتی به مشکل می‌خورد، سراغِ تنظیماتِ خاصِ HotSpot می‌رود، نه اینکه فکر کند این ذاتِ زبان است.


بخش ۲ — بارگذاری کلاس: loading → linking → initialization

وقتی برنامه‌ات اجرا می‌شود، JVM همهٔ کلاس‌ها را یک‌جا لود نمی‌کند. فقط وقتی یک کلاس را لود می‌کند که برای اولین بار واقعاً به آن نیاز پیدا کنی. به این می‌گویند بارگذاری تنبل (lazy loading).

دایرة‌المعارف ۱۰ جلدی

وقتی یک دایرة‌المعارف ۱۰ جلدی می‌خری، همان لحظه هر ۱۰ جلد را نمی‌خوانی! فقط وقتی جلدِ «حرف ک» را باز می‌کنی که واقعاً کلمه‌ای با «ک» لازم داشته باشی. JVM هم همین‌طور است: کلاس را دقیقاً لحظه‌ای لود می‌کند که اولین بار به آن برسی (مثلاً وقتی new می‌زنی یا یک متد static صدا می‌کنی). این باعث می‌شود برنامه سریع‌تر بالا بیاید و حافظهٔ الکی اشغال نشود.

وقتی JVM بالاخره به یک کلاس نیاز پیدا کرد، آن را طبق استاندارد (JLS §12.4 و JVMS §5) در سه فاز با ترتیب دقیق آماده می‌کند. بیا با تشبیهِ استخدام یک کارمند جدید ببینیمشان.

فاز ۱: بارگذاری (Loading)

بخش «منابع انسانی» (همان ClassLoader) بایت‌های فایل .class را (از روی دیسک، از داخل یک JAR، از شبکه، یا حتی بایت‌های تولیدشده در لحظه) می‌خواند و یک پروندهٔ پرسنلی در حافظه می‌سازد — در جاوا این پرونده یک آبجکت از نوع Class<?> است که در heap قرار می‌گیرد.

هویت یک کلاس فقط «اسمش» نیست

هویتِ زمان‌اجرای یک کلاس برابر است با ترکیبِ (نامِ کاملِ کلاس + همان ClassLoaderی که لودش کرده). یعنی اگر دقیقاً همان بایت‌ها توسط دو ClassLoader مختلف لود شوند، JVM آن‌ها را دو نوعِ کاملاً متفاوت می‌بیند!

خطای گیج‌کنندهٔ `X cannot be cast to X`

به‌خاطر همان قانون بالا، ممکن است خطای عجیبی ببینی: ClassCastException: com.x.Foo cannot be cast to com.x.Foo — یعنی «Foo را نمی‌توان به Foo تبدیل کرد»! چطور ممکن است؟ چون این دو Foo را دو ClassLoader مختلف لود کرده‌اند و از نظر JVM دو نوعِ جدا هستند. این تله در سرورهای اپلیکیشن، OSGi و hot-reload خیلی رخ می‌دهد.

فاز ۲: لینک (Linking) — سه زیرمرحله

حالا که پرونده ساخته شد، باید کارمند را به سیستمِ شرکت وصل کنیم. سه زیرمرحله دارد:

  1. وارسی (Verification) — بازرسی امنیتی: بایت‌کد از نظر امنیت و درستی چک می‌شود: آیا type-safe است؟ آیا عملیاتِ روی پشته معتبرند؟ آیا پرشِ غیرمجاز به جایی از کد ندارد؟ این ستون فقراتِ امنیتِ JVM است — همین وارسی است که نمی‌گذارد یک بایت‌کدِ دستکاری‌شده حافظه را خراب کند.

  2. آماده‌سازی (Preparation) — میز خالی: فیلدهای static ساخته می‌شوند و مقدارِ پیش‌فرض می‌گیرند (عددها 0، آبجکت‌ها null، بولین‌ها false) — نه مقدار واقعی. مثل اینکه به کارمند یک میز خالی و دفترچهٔ سفید می‌دهی، نه ابزار کارِ واقعی.

  3. حل ارجاع (Resolution) — تبدیل وعده به آدرس: در کد نوشته‌ای «برو با کلاس Database کار کن». این فعلاً یک نامِ نمادین (symbolic) و مبهم است. اینجا JVM آن نام را به ارجاعِ مستقیم (آدرسِ دقیقِ آن چیز در حافظه) تبدیل می‌کند. این مرحله هم می‌تواند تنبل باشد و تا لحظهٔ آخر عقب بیفتد.

فاز ۳: مقداردهی اولیه (Initialization)

حالا «روز اول کاری». JVM یک متدِ مخفیِ خاص به اسم <clinit> (مخففِ class initializer) را اجرا می‌کند. این متد را خودت نمی‌نویسی؛ کامپایلر آن را از دو چیز می‌سازد و از بالا به پایین اجرا می‌کند:

  • مقداردهیِ واقعیِ فیلدهای static (همان 0ها و nullهای مرحلهٔ قبل حالا مقدار واقعی می‌گیرند)،
  • و بلوک‌های static { ... }.

کِی اجرا می‌شود؟ فقط با «استفادهٔ فعال» از کلاس: با new، فراخوانی متد static، خواندن/نوشتنِ فیلد static (مگر یک ثابتِ زمانِ کامپایل باشد)، reflection، یا مقداردهیِ یک زیرکلاس.

تضمینِ طلاییِ JVM

‏JVM تضمین می‌کند <clinit> دقیقاً یک‌بار و به‌صورتِ کاملاً thread-safe اجرا شود. یعنی اگر ۱۰۰ ترد هم‌زمان اولین بار به این کلاس برسند، JVM خودش صف می‌بندد: فقط یکی مقداردهی می‌کند و بقیه منتظر می‌مانند تا آماده شود. بدون اینکه تو حتی یک کلمه synchronized بنویسی.

از همین تضمین، یک الگوی بسیار محبوب و هوشمند برای Singletonِ تنبل ساخته می‌شود:

public class Config {
    private Config() {}                       // کسی از بیرون نمی‌تواند بسازدش

    // این کلاس داخلی تا وقتی getInstance() صدا نخورد، اصلاً لود نمی‌شود (بارگذاری تنبل).
    private static class Holder {
        static final Config INSTANCE = new Config(); // <clinit> یک‌بار و thread-safe توسط JVM
    }

    public static Config getInstance() { return Holder.INSTANCE; }
}
چرا این الگو نابغه‌وار است

تا وقتی کسی getInstance() را صدا نزند، کلاسِ Holder اصلاً وجودِ خارجی پیدا نمی‌کند (تنبل). لحظه‌ای که صدا زده شد، JVM Holder را لود می‌کند و چون خودش تضمین می‌کند <clinit> فقط یک‌بار و امن اجرا شود، ما یک Singletonِ تنبل، امن در برابر چند-تردی، و بدونِ قفلِ دستی ساختیم — بهترینِ هر دو دنیا.

سلسله‌مراتب ClassLoader و مدل واگذاری (Delegation)

بخش‌های منابع انسانی (ClassLoaderها) به‌صورت زنجیرهٔ والد-فرزند کار می‌کنند و یک قانونِ سفت‌وسخت دارند: «اول از رئیست بپرس.» به این می‌گویند مدل واگذاری به والد (parent-delegation): هر loader اول از والدش می‌خواهد کلاس را لود کند؛ فقط اگر والد نتوانست، خودش دست‌به‌کار می‌شود.

Bootstrap ClassLoader   (بومی/C++، هستهٔ جاوا را لود می‌کند: java.lang.* و ...)
   └─ Platform ClassLoader   (قبلاً «extension»؛ ماژول‌های استانداردِ خودِ JDK)
        └─ System/Application ClassLoader   (classpathِ برنامهٔ تو)
             └─ ClassLoaderهای سفارشی / وب‌اپ‌ها (مثلاً Tomcat)

چرا این قانون وجود دارد؟ امنیت. بیا با یک مثال ببینیم:

مهرِ رسمیِ شرکت

فرض کن مدیرعامل (Bootstrap) مهرِ رسمی و اصلیِ شرکت را در گاوصندوقش دارد — معتبرترین مهرِ دنیا. یک هکر یک مهرِ تقلبیِ دقیقاً شبیهِ آن می‌سازد و یواشکی روی میزِ تو (پوشهٔ پروژه‌ات) می‌گذارد؛ اسمش را هم می‌گذارد java.lang.String.

اگر قانونِ واگذاری نبود: تو نیاز به مهر داری، به میزِ خودت نگاه می‌کنی، مهرِ تقلبی را برمی‌داری و اسناد را با آن مهر می‌کنی. هکر برنده شد — کدِ Stringِ ویروسیِ او به‌جای Stringِ امنِ جاوا اجرا می‌شود و کل سیستم آلوده می‌شود.

با قانونِ واگذاری: تو حق نداری سرخود از میزت مهر برداری. اول از مدیرعامل (Bootstrap) می‌پرسی. او می‌گوید «خودم نسخهٔ اصلیِ String را دارم، بیا این را استفاده کن.» پس مهرِ تقلبیِ روی میزت اصلاً دیده نمی‌شود. هکر شکست خورد.

چرا واگذاری امنیت می‌آورد

واگذاری مانع می‌شود کسی با گذاشتنِ یک فایلِ هم‌نام با کلاس‌های هستهٔ جاوا (مثل java.lang.String) در پوشهٔ پروژه، نسخهٔ اصلی را سایه بزند (shadow کند). چون JVM همیشه اول از هسته و والدهای بالادست می‌پرسد، نسخهٔ اصلی و امن همیشه برنده است.

دو نکتهٔ نسخه‌ای که خوب است بدانی

از Java 9 به بعد، loaderِ قدیمیِ «extension» به Platform ClassLoader تبدیل شد و کل ماجرا روی سیستمِ ماژول (JPMS) بنا شد. ضمناً چون Bootstrap با C++ نوشته شده و یک آبجکتِ جاوایی نیست، اگر SomeCoreClass.class.getClassLoader() را صدا بزنی، null می‌گیری (یعنی «من را Bootstrap لود کرده»).

استثنای مهم: سرورهای وب مثل Tomcat قانون را برعکس می‌کنند

در یک Tomcat ممکن است ۱۰ برنامهٔ وبِ مختلف هم‌زمان اجرا شوند و هر کدام نسخهٔ متفاوتی از یک کتابخانه (مثلاً Log4j) بخواهند. اگر همه از والد بپرسند، تداخل می‌شود. برای همین Tomcat عمداً به loaderِ هر وب‌اپ می‌گوید: «اول خودت پوشه‌ات را بگرد؛ اگر پیدا نکردی، بعد از والد بپرس.» این معکوسِ واگذاریِ استاندارد است و هدفش ایزوله‌کردنِ برنامه‌هاست تا همدیگر را خراب نکنند.


بخش ۳ — نواحی دادهٔ زمان‌اجرا: حافظه کجا زندگی می‌کند؟

این بخش به تو نشان می‌دهد وقتی برنامه اجرا می‌شود، هر چیزی کجا ذخیره می‌شود. بیا با تشبیهِ آشپزخانه شروع کنیم و بعد جدولِ دقیق را ببینیم.

آشپزخانهٔ رستوران
  • Heap (هیپ) = یخچالِ بزرگ: هر چیزی که با new می‌سازی (همهٔ آبجکت‌ها و آرایه‌ها) اینجا با حجمِ زیاد نگهداری می‌شود. این ناحیه توسط Garbage Collector مدیریت می‌شود.
  • Stack (پشته) = میزِ کارِ لحظه‌ایِ سرآشپز: چیزهای کوچک و موقتیِ هر ترد (متغیرهای محلی، ارجاع‌ها، آدرسِ برگشت) اینجا می‌آیند و با تمام‌شدنِ متد پاک می‌شوند.
  • Metaspace = کتابخانهٔ دستورپخت‌ها: اطلاعاتِ خودِ کلاس‌ها (نه آبجکت‌ها) اینجاست — ساختار، متدها، بایت‌کد.

حالا جدولِ کاملِ نواحی و اینکه هر کدام اگر پر شود چه خطایی می‌دهد:

ناحیه مشترک بین کیست؟ چه چیزی نگه می‌دارد خطای پرشدن
Heap کلِ JVM همهٔ آبجکت‌ها و آرایه‌ها OutOfMemoryError: Java heap space
Metaspace کلِ JVM متادیتای کلاس، بایت‌کدِ متدها OutOfMemoryError: Metaspace
JVM Stack هر ترد جدا فریم‌ها (متغیرهای محلی، operandها، آدرس برگشت) StackOverflowError
PC Register هر ترد جدا آدرسِ دستورِ بایت‌کدِ جاری
Native Method Stack هر ترد جدا فریم‌های کدِ C/C++ (از طریق JNI)
Code Cache کلِ JVM کدِ ماشینِ کامپایل‌شده توسط JIT CodeCache is full (JIT خاموش می‌شود)

دو نکتهٔ مهم دربارهٔ این جدول:

۱. هر ترد پشتهٔ خودش را دارد. اگر یک متد بی‌نهایت خودش را صدا بزند (بازگشتِ بی‌پایان)، پشته پر می‌شود و StackOverflowError می‌گیری. از طرف دیگر، اگر ده‌ها هزار تردِ سیستمی (platform thread) بسازی، حافظهٔ بومیِ سیستم تمام می‌شود (هر پشته حدود ۵۱۲KB تا ۱MB است، قابل‌تنظیم با -Xss).

چرا Virtual Thread ها (جاوا ۲۱) اهمیت دارند

همین فشارِ «هر ترد یک پشتهٔ بومیِ گران می‌خواهد» است که virtual threadها حلش می‌کنند: آن‌ها پشتهٔ بومی را به‌صورت دائم اشغال (pin) نمی‌کنند، پس می‌توانی میلیون‌ها تای‌شان را داشته باشی. (فصلِ مخصوصِ خودش را دارد.)

۲. Metaspace جای PermGen را گرفت. این یکی از مهم‌ترین تغییرات جاوا ۸ است و سؤالِ رایجِ مصاحبه — در بخش بعد کاملاً بازش می‌کنیم.

تمرکز روی Metaspace: چرا و چطور جای PermGen را گرفت

کارخانهٔ خودروسازی — تفاوتِ Heap و Metaspace

یک کارخانهٔ خودروسازی را تصور کن:

  • Heap = پارکینگِ بزرگ: هر خودرویی که ساخته می‌شود (مثلاً ۱۰٬۰۰۰ خودروی یک‌مدل) اینجا پارک می‌شود. در جاوا، این خودروها همان آبجکت‌ها هستند.
  • Metaspace = اتاقِ نقشه و مهندسی: جایی که نقشهٔ ساختِ (blueprint) خودرو نگهداری می‌شود.

نکتهٔ کلیدی: برای ساختِ ۱۰٬۰۰۰ خودرو، فقط به یک نقشه نیاز داری، نه ۱۰٬۰۰۰ نقشه! در جاوا هم دقیقاً همین است — هزاران آبجکت از یک کلاس می‌سازی (همه در Heap)، اما اطلاعاتِ ساختاریِ خودِ کلاس فقط یک بار در Metaspace ذخیره می‌شود.

دقیقاً چه چیزی در Metaspace است؟ متادیتا (یعنی «داده دربارهٔ داده»): نامِ کلاس، لیستِ متدها و پارامترهایشان، لیستِ فیلدها و نوعِشان، اطلاعاتِ ارث‌بری، و Constant Pool. توجه: مقدارِ واقعیِ یک فیلد (مثلاً اینکه نامِ کاربر «علی» است) در Heap است؛ اما این واقعیت که «کلاسِ User یک فیلدِ name از نوع String دارد» در Metaspace است.

چه چیزی از PermGen به Metaspace تغییر کرد؟

قبل از جاوا ۸ این فضا PermGen نام داشت و داخلِ Heap بود با یک سقفِ ثابت و کوچک. اگر برنامه کلاس‌های زیادی لود می‌کرد، پر می‌شد و خطای معروفِ OutOfMemoryError: PermGen space می‌گرفتی. از جاوا ۸، این فضا به حافظهٔ بومیِ سیستم‌عامل (off-heap) منتقل شد و اسمش شد Metaspace. مزیت: حالا پویا رشد می‌کند و تا وقتی رمِ فیزیکی پر نشده، با یک سقفِ کوچکِ ثابت کرش نمی‌کند.

پس چرا هنوز `OutOfMemoryError: Metaspace` می‌بینیم؟

دو علتِ اصلی: (۱) نشتیِ ClassLoader — مخصوصاً در سرورهای وب: هر بار که کد را redeploy می‌کنی، Tomcat یک ClassLoaderِ جدید می‌سازد و همهٔ کلاس‌ها را دوباره لود می‌کند؛ اگر ClassLoaderهای قدیمی درست دور ریخته نشوند، نسخه‌های تکراری در Metaspace انباشته می‌شوند. (۲) تولیدِ کلاس‌های داینامیک — کتابخانه‌هایی مثل Hibernate/Spring/CGLIB در زمان اجرا کلاس می‌سازند؛ اگر افسارگسیخته شوند، Metaspace را پر می‌کنند. برای همین در پروداکشن با -XX:MaxMetaspaceSize سقف می‌گذاری.


بخش ۴ — چیدمانِ آبجکت و هدرها (HotSpot، ۶۴ بیتی)

حالا که می‌دانیم آبجکت‌ها در Heap زندگی می‌کنند، ببینیم یک آبجکت دقیقاً چه شکلی در حافظه است. هر آبجکتِ heap علاوه بر فیلدهای خودت، یک هدر (header) دارد که JVM برای مدیریتش لازم دارد:

[ mark word: 8 بایت ][ klass pointer: 4 بایت ][ فیلدهای تو... ][ padding تا مضربِ ۸ ]
  • Mark word — چیزهای مدیریتی: hashCodeِ هویتی، بیت‌های سن (age) برای GC، وضعیتِ قفل، و هنگام جابه‌جایی در GC اشاره‌گرِ forwarding.
  • Klass pointer — به متادیتای کلاس در Metaspace اشاره می‌کند (یعنی «من از نوعِ کدام کلاسم؟»).
compressed oops چیست و چرا اشاره‌گر ۴ بایت است

وقتی heap کوچک‌تر از ~۳۲GB است (پیش‌فرض)، HotSpot از compressed oops استفاده می‌کند: اشاره‌گرها را به‌جای ۸ بایت در ۴ بایت ذخیره می‌کند (به‌صورتِ آفستِ مقیاس‌دار). این یک صرفه‌جوییِ بزرگِ حافظه است — به همین برمی‌گردیم در بخشِ «چرا ۳۱GB بهتر از ۳۳GB است».

پس یک Objectِ خالی ۱۶ بایت می‌شود (۱۲ بایت هدر + ۴ بایت padding). آرایه‌ها یک فیلدِ ۴ بایتیِ «طول» هم اضافه دارند.

چرا این برای کارایی مهم است

همین سربارِ هدر توضیح می‌دهد چرا boxing گران است: یک Integerِ باکس‌شده حدود ۱۶ بایت (هدر + مقدار + padding) به‌علاوهٔ یک ارجاعِ جداگانه می‌گیرد، در حالی که یک intِ خام فقط ۴ بایت است. در حلقه‌های داغ و ساختمان‌دادهٔ بزرگ، این تفاوت واقعاً حس می‌شود.

نگاه به آینده

‏Java 24 با JEP 450 (هدرهای فشردهٔ آبجکت، فعلاً آزمایشی) هدر را به ۸ بایت می‌رساند. دانستنش خوب است، اما هنوز پیش‌فرض نیست.


بخش ۵ — مبانیِ Garbage Collection

اینجا می‌رسیم به جادوی جاوا: تو حافظه را دستی آزاد نمی‌کنی؛ یک زباله‌جمع‌کن (Garbage Collector) خودکار این کار را می‌کند. اما چطور می‌فهمد کدام آبجکت «زباله» است؟

GC بر پایهٔ «قابلیت دسترسی» کار می‌کند، نه شمارش

نخِ متصل به لنگر

تصور کن هر آبجکت یک بادکنک است و ارجاع‌ها نخ‌هایی که بادکنک‌ها را به هم و به یک لنگرِ ثابت (GC root) وصل می‌کنند. لنگرها چیزهایی هستند که همیشه «زنده»‌اند: متغیرهای محلیِ تردهای درحال‌اجرا، فیلدهای static، و ارجاع‌های JNI.

‏GC از لنگرها شروع می‌کند و هر بادکنکی را که با نخ‌ها بتوان به یک لنگر رسید، «زنده» علامت می‌زند (mark). هر بادکنکی که هیچ مسیری به لنگر نداشته باشد، زباله است و آزاد می‌شود.

چرا این روش از «شمارشِ ارجاع» بهتر است

چون چرخه‌ها را درست مدیریت می‌کند. فرض کن آبجکت A به B اشاره کند و B به A، ولی هیچ‌کدام به یک لنگر وصل نباشند — یک «جزیرهٔ جدا». در روشِ سادهٔ شمارشِ ارجاع، شمارندهٔ هر دو ۱ است پس هرگز آزاد نمی‌شوند (نشتی!). اما GCِ ردیابی‌محورِ (tracing) جاوا چون از لنگر شروع می‌کند و به این جزیره نمی‌رسد، هر دو را درست تشخیصِ مرده می‌دهد.

فرضیهٔ نسلی (Generational Hypothesis)

یک مشاهدهٔ تجربی که کلِ طراحیِ GCهای مدرن رویش بنا شده: بیشترِ آبجکت‌ها جوان می‌میرند. یعنی اکثرِ آبجکت‌ها خیلی زود پس از ساخته‌شدن بی‌مصرف می‌شوند (مثل متغیرهای موقتِ داخلِ یک متد). GC از این حقیقت با تقسیمِ heap بهره می‌برد:

نسلِ جوان (Young):  [ Eden | Survivor S0 | Survivor S1 ]      نسلِ قدیم (Old / Tenured)
  • آبجکت‌های تازه در Eden ساخته می‌شوند (تخصیصِ فوق‌سریع، فقط با جابه‌جاییِ یک اشاره‌گر).
  • یک GC جوان (minor GC) آبجکت‌های زندهٔ Eden و یک Survivor را به Survivor دیگر کپی می‌کند و سن‌شان را یکی زیاد می‌کند. چون فقط آبجکت‌های زنده را لمس می‌کند و بیشترشان مرده‌اند، خیلی ارزان است.
  • آبجکتی که به‌اندازهٔ کافی چرخهٔ جوان را رد کند (-XX:MaxTenuringThreshold، تا ۱۵)، به نسلِ قدیم (Old) «ترفیع (promote/tenure)» می‌یابد.
  • یک GC قدیم (major GC) نسلِ قدیم را جمع می‌کند؛ و یک Full GC همه‌چیز را (جوان + قدیم + اغلب metaspace). Full GC همان موردِ گران‌قیمتی است که باید ازش پرهیز کنی.
Stop-the-World یعنی چه

برای اینکه GC بتواند بی‌خطر آبجکت‌ها را جابه‌جا کند و لنگرها را بشمارد، JVM لحظه‌ای همهٔ تردهای برنامه را در یک نقطهٔ امن (safepoint) متوقف می‌کند — به این توقف می‌گویند Stop-the-World (STW). هر GCی مقداری STW دارد. GCهای مدرن این توقف را با انجامِ بیشترِ کار به‌صورتِ هم‌زمان (concurrent) با برنامه کمینه می‌کنند. سیستم‌های حساس به تأخیر، زندگی و مرگشان با طول و تعدادِ همین توقف‌هاست.

ارجاع‌های ضعیف: weak / soft / phantom

جاوا چند نوع ارجاعِ خاص دارد که به تو اجازه می‌دهند به GC بگویی «این آبجکت آن‌قدرها هم مهم نیست»:

  • SoftReference — فقط زیرِ فشارِ حافظه پاک می‌شود. مناسبِ کش (ولی مراقب باش، می‌تواند نشتی بدهد).
  • WeakReference — در اولین GC بعد از اینکه فقط ضعیف قابل‌دسترس شد پاک می‌شود. مبنای WeakHashMap.
  • PhantomReference — برای پاک‌سازیِ قطعیِ منابعِ بومی بعد از مرگِ آبجکت (کلاسِ Cleaner، جایگزینِ مدرنِ finalize()).

بخش ۶ — چشم‌انداز GCها: کدام کِی؟

جاوا چند Garbage Collector مختلف دارد و می‌توانی با یک فلگ انتخابشان کنی. هر کدام یک مصالحه (trade-off) بین throughput (کارِ کلیِ انجام‌شده) و latency (کوتاهیِ توقف‌ها) دارند.

Collector فلگ مدلِ توقف بهترین برای مصالحه
Serial -XX:+UseSerialGC کاملاً STW، تک‌رشته‌ای heapِ کوچک، کانتینرِ تک‌CPU، ابزارِ CLI ساده‌ترین و کم‌سربارترین؛ ولی در مقیاسِ بزرگ توقف‌های طولانی
Parallel -XX:+UseParallelGC کاملاً STW، چندرشته‌ای jobهای batch، بیشینه‌کردنِ throughput بهترین throughput؛ بدترین tail-latency
G1 -XX:+UseG1GC (پیش‌فرض از Java 9) مارکِ عمدتاً concurrent، تخلیهٔ STW همه‌منظوره، heapِ بزرگ، تعادلِ latency/throughput هدفِ زمانِ توقف می‌دهی (-XX:MaxGCPauseMillis، پیش‌فرض ۲۰۰ms)
ZGC -XX:+UseZGC concurrent، توقفِ زیرِ ۱ms heapِ بسیار بزرگ (تا ترابایت)، SLAی سختِ تأخیر throughputِ کمی کمتر، سربارِ CPU/حافظهٔ بیشتر
Shenandoah -XX:+UseShenandoahGC concurrent، توقفِ زیرِ ۱ms تأخیرِ پایین، بیلدهای Red Hat مشابهِ ZGC
پیش‌فرض چیست؟ (سؤالِ رایجِ مصاحبه)

در همهٔ JDKهای مدرن — از جمله Java 17 و Java 21 — Collectorِ پیش‌فرض همچنان G1 است. ZGC و Shenandoah اختیاری (opt-in) هستند. اگر کسی بپرسد «پیش‌فرضِ ۱۷ و ۲۱؟»، جوابِ درست: G1 (که از Java 9 جایگزینِ Parallel شد).

G1 با کمی جزئیات

‏G1 (مخففِ Garbage-First) هیپ را به حدودِ ۲۰۴۸ ناحیهٔ (region) مساوی تقسیم می‌کند (هر کدام ۱ تا ۳۲MB). هر ناحیه به‌صورتِ پویا برچسبِ Eden، Survivor، Old یا Humongous (آبجکت‌های خیلی بزرگ‌تر از نصفِ یک ناحیه) می‌گیرد. G1 عمدتاً به‌صورتِ concurrent مارک می‌کند، بعد توقف‌های تخلیه (evacuation) ای انجام می‌دهد که آبجکت‌های زنده را اول از ناحیه‌هایی بیرون می‌کشد که بیشترین زباله را دارند — برای همین اسمش «زباله‌اول» است. تو یک هدفِ توقف می‌دهی، نه یک چیدمانِ دقیق؛ G1 خودش تصمیم می‌گیرد چند ناحیه را جمع کند تا به هدفت برسد. G1 عملاً Full GCهای چندثانیه‌ایِ collectorِ قدیمیِ CMS را حذف کرد (CMS در Java 14 کاملاً برداشته شد).

ZGC و Generational ZGC — نسخه‌ها را دقیق بگو

‏ZGC یک collectorِ concurrent، ناحیه‌محور و فشرده‌ساز است که با ترفندهایی به اسمِ اشاره‌گرهای رنگی (colored pointers) و load barrier آبجکت‌ها را حین اجرای برنامه جابه‌جا می‌کند. تیترِ اصلی‌اش: توقف‌ها زیرِ یک میلی‌ثانیه‌اند (معمولاً ۰٫۰۵ تا ۰٫۵ms) و با بزرگ‌شدنِ heap رشد نمی‌کنند — تا heapهای چند-ترابایتی مقیاس می‌پذیرد.

تاریخچهٔ نسخه‌های ZGC (این را دقیق حفظ کن)
  • Java 15: ZGC آمادهٔ تولید شد (JEP 377).
  • Java 21: نسخهٔ Generational ZGC اضافه شد (JEP 439)، ولی اختیاری بود: -XX:+UseZGC -XX:+ZGenerational. صرفِ -XX:+UseZGC نسخهٔ قدیمیِ غیرنسلی را می‌داد.
  • Java 23: حالتِ نسلی برای ZGC پیش‌فرض شد (JEP 474) و فلگِ ZGenerational منسوخ اعلام شد.
  • Java 24: حالتِ غیرنسلی کاملاً حذف شد (JEP 490).

دقیق‌گفتنِ این نسخه‌ها به مصاحبه‌گر نشان می‌دهد پلتفرم را واقعاً دنبال می‌کنی.

کِی ZGC به‌جای G1؟

وقتی heap بزرگ است و نیازِ سختی به کوتاهیِ tail-latency داری (سیستم‌های ترید، سرویس‌های فوق‌کم‌تأخیر) که حتی توقف‌های ~۱۰۰-۲۰۰ میلی‌ثانیه‌ایِ G1 هم برایت زیاد است. برای سرویس‌های عمومی روی G1 بمان — معمولاً throughputِ بهتر و سربارِ حافظهٔ کمتری دارد و پیش‌فرضِ آزموده است.


بخش ۷ — کامپایلِ JIT، tiered و inlining

یادت هست گفتیم JVM اول کد را تفسیر می‌کند و بخش‌های داغ را کامپایل؟ حالا دقیق‌ترش می‌کنیم.

HotSpot ابتدا بایت‌کد را تفسیر (interpret) می‌کند و هم‌زمان از آن پروفایل می‌گیرد: هر متد چند بار صدا خورده؟ کدام branchها بیشتر گرفته می‌شوند؟ کدام تایپ‌ها واقعاً می‌آیند؟ متدهای داغ توسط دو کامپایلر به کدِ ماشین تبدیل می‌شوند:

  • C1 (client): کامپایلِ سریع، بهینه‌سازیِ سبک → استارتاپِ سریع.
  • C2 (server): کامپایلِ کند، بهینه‌سازیِ تهاجمی (inlining، بازکردنِ حلقه، حذفِ کدِ مرده) → اوجِ کارایی.
آشپزِ مبتدی در برابر سرآشپزِ نابغه

مفسر مثل یک آشپزِ مبتدی است: بایت‌کد را خط‌به‌خط می‌خواند و اجرا می‌کند (سرعتِ معمولی، ولی همین الان شروع می‌کند). JIT مثل سرآشپزی است که حواسش هست کدام غذا صد بار سفارش داده شده؛ آن دستور را از قبل به «حرکاتِ عضلانیِ خودکار» (کدِ ماشینِ بومی) تبدیل می‌کند تا دفعهٔ بعد با سرعتِ نور بپزد.

Tiered compilation (پیش‌فرض) هر دو کامپایلر را در ۵ سطح ترکیب می‌کند:

Level 0: مفسر (interpreter)
Level 1: C1، بدونِ پروفایل (متدهای بدیهی)
Level 2: C1، پروفایلِ محدود
Level 3: C1، پروفایلِ کامل      ← بیشترِ متدها اینجا گرم می‌شوند
Level 4: C2، کاملاً بهینه        ← داغ‌ترین متدها به اینجا می‌رسند

کد از تفسیر شروع می‌شود، بعد از گرم‌شدن با پروفایل توسط C1 کامپایل می‌شود، و مسیرهای واقعاً داغ به C2 «فارغ‌التحصیل» می‌شوند. کدِ کامپایل‌شده در code cache زندگی می‌کند.

حادثهٔ واقعیِ پروداکشن: پرشدنِ Code Cache

اگر code cache پر شود، JIT خاموش می‌شود و همه‌چیز به مفسرِ کُند برمی‌گردد — یعنی افتِ ناگهانیِ کارایی. لاگش CodeCache is full است. در برنامه‌های خیلی بزرگ گاهی باید سقفش را زیاد کنی (-XX:ReservedCodeCacheSize).

Inlining — باارزش‌ترین بهینه‌سازی

Inlining یعنی جایگزین‌کردنِ فراخوانیِ یک متد با بدنهٔ خودِ آن متد. چرا مهم است؟ چون وقتی بدنه سرِ جایش کپی شد، همهٔ بهینه‌سازی‌های دیگر می‌توانند از مرزِ متد عبور کنند و با هم ترکیب شوند. C2 متدهای کوچک و داغ را تهاجمی inline می‌کند.

چرا «یک getter اضافه کن» در جاوا رایگان است

خیلی‌ها می‌ترسند getter/setter اضافه کنند چون فکر می‌کنند «هزینهٔ فراخوانیِ متد» دارد. در عمل C2 این متدهای کوچک را inline می‌کند و کاملاً محو می‌شوند — انگار مستقیم به فیلد دست زده باشی. پس نگرانِ کاراییِ getterها نباش.

تحلیلِ فرار (Escape Analysis) و جایگزینیِ اسکالر

‏C2 سعی می‌کند اثبات کند آیا یک آبجکت از متد/تردِ سازنده‌اش فرار می‌کند (یعنی جایی بیرون هم دیده می‌شود) یا نه. اگر ثابت شود هرگز فرار نمی‌کند:

  • جایگزینیِ اسکالر (scalar replacement): آبجکت اصلاً روی heap ساخته نمی‌شود؛ فیلدهایش مستقیم در رجیستر/پشته می‌نشینند. برای همین یک حلقهٔ داغ که یک Pointِ موقت می‌سازد می‌تواند صفر زباله تولید کند.
  • حذفِ قفل (lock elision): اگر روی آبجکتی که فرار نمی‌کند synchronized باشد، آن قفل کلاً حذف می‌شود.
// C2 می‌تواند اثبات کند p فرار نمی‌کند؛ پس آبجکتِ موقت شاید هرگز به heap نرسد.
int sumOfSquares(int[] xs, int[] ys) {
    int total = 0;
    for (int i = 0; i < xs.length; i++) {
        Point p = new Point(xs[i], ys[i]); // نامزدِ جایگزینیِ اسکالر
        total += p.x * p.x + p.y * p.y;
    }
    return total;
}
روی تحلیلِ فرار حساب نکن

تحلیلِ فرار تضمین نیست — بهترین‌تلاش است و می‌تواند شکست بخورد (مثلاً وقتی آبجکت به یک متدِ inline‌نشده پاس شود). درستیِ برنامه‌ات را بر پایهٔ آن طراحی نکن. و اگر داری با JMH بنچمارک می‌گیری، بدان که به همین دلیل JMH از Blackhole استفاده می‌کند تا نگذارد کامپایلر کدِ «به‌ظاهر بی‌مصرفِ» تو را کلاً حذف کند.

دی‌اپتیمایزیشن (Deoptimization)

‏C2 شرط‌بندی‌های حدسی می‌کند (مثلاً «این نقطهٔ فراخوانی همیشه فقط ArrayList دیده، پس فرض می‌کنم همیشه همین است و کدِ بهینه می‌سازم»). اگر یک روز واقعیت این فرض را نقض کند (ناگهان یک LinkedList بیاید)، JVM آن کدِ کامپایل‌شده را دور می‌ریزد و موقتاً به مفسر برمی‌گردد و شاید دوباره کامپایل کند. این به‌صورتِ یک افتِ کاراییِ کوتاه بعد از ظهورِ یک تایپِ جدید دیده می‌شود.


بخش ۸ — فلگ‌های کلیدی که هر سنیور باید بداند

# اندازهٔ heap — در پروداکشن min == max بگذار تا از توقفِ resize و تکه‌تکه‌شدن جلوگیری شود
-Xms4g -Xmx4g

# کانتینر-آگاه (از Java 10+ پیش‌فرض فعال): heap را درصدی از حافظهٔ کانتینر بگیر
-XX:MaxRAMPercentage=75.0

# انتخابِ collector
-XX:+UseG1GC            # پیش‌فرض
-XX:+UseZGC             # از Java 23+ به‌صورتِ نسلی
-XX:MaxGCPauseMillis=100

# اندازهٔ پشتهٔ هر ترد
-Xss512k

# سقفِ Metaspace (وگرنه تا اتمامِ حافظهٔ بومی رشد می‌کند)
-XX:MaxMetaspaceSize=256m

# هنگام OOM: heap dump بگیر و سریع بمیر (برای عیب‌یابیِ بعدی)
-XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/dumps
-XX:+ExitOnOutOfMemoryError

# لاگِ GC (یکپارچه، از Java 9+)
-Xlog:gc*:file=gc.log:time,uptime,level,tags:filecount=5,filesize=20m
چرا `-Xms == -Xmx`؟

اگر min و max یکی نباشند، heapِ متعهدشده مدام کوچک و بزرگ می‌شود که خودش باعثِ Full GC و page fault می‌گردد. با ثابت‌کردنِ اندازه، این نوسان حذف می‌شود. در کانتینرها MaxRAMPercentage را ترجیح بده تا JVM به محدودیتِ cgroup احترام بگذارد، نه اینکه کلِ رمِ هاست را ببیند و بعد OOM-kill شود.


بخش ۹ — نشتیِ حافظه و طبقه‌بندیِ OutOfMemoryError

«نشتیِ حافظه» در جاوا یعنی چه؟

در زبانی که GC دارد، نشتیِ حافظه یعنی قابلیتِ دسترسیِ ناخواسته: آبجکت‌هایی که دیگر کارت با آن‌ها تمام شده، ولی هنوز از یک لنگرِ GC قابل‌دسترس‌اند، پس GC حق ندارد آزادشان کند و روی هم انباشته می‌شوند.

منابعِ کلاسیکِ نشتی:

  • کالکشن‌های static که فقط رشد می‌کنند (static Map cache = ... که هرگز چیزی از آن حذف نمی‌شود).
  • کش‌های بی‌کران — به‌جایش از سقفِ اندازه استفاده کن (Caffeine، LRU).
  • listenerها/callbackهایی که هرگز unregister نمی‌شوند — سوژه، observer را برای همیشه نگه می‌دارد.
  • ThreadLocal در thread poolها — تردِ pool‌شده هرگز نمی‌میرد، پس مقدارِ ThreadLocal هم هرگز آزاد نمی‌شود؛ همیشه در finally آن را remove() کن.
  • نشتیِ ClassLoader — یک ارجاعِ باقی‌مانده به کلاس‌های وب‌اپ، کلِ ClassLoader (و همهٔ کلاس‌هایش در Metaspace) را هنگامِ redeploy زنده نگه می‌دارد.

انواعِ OutOfMemoryError (هر کدام معنای متفاوتی دارد)

پیام معنا علتِ معمول
Java heap space heap پر شد و GC نتوانست کافی آزاد کند نشتی یا heapِ کوچک
GC overhead limit exceeded بیش از ۹۸٪ زمان صرفِ GC می‌شود، کمتر از ۲٪ آزاد می‌شود heap تقریباً پر، thrashing
Metaspace فضای متادیتای کلاس تمام شد نشتیِ classloader، کلاس‌های داینامیکِ زیاد
unable to create new native thread حافظهٔ بومیِ OS یا محدودیتِ ulimit پر شد نشتیِ ترد، -Xssِ بزرگ × تردهای زیاد
Direct buffer memory ByteBufferِ مستقیمِ off-heap تمام شد بافرهای مستقیمِ نشتی
Requested array size exceeds VM limit آرایه‌ای نزدیکِ Integer.MAX_VALUE خواستی باگِ منطقی
نکتهٔ مهم: `OutOfMemoryError` را `catch` نکن

این یک Error است، نه Exception. گرفتنش تقریباً همیشه اشتباه است، چون JVM ممکن است در وضعیتِ غیرقابلِ‌بازیابی باشد. بگذار برنامه سریع بمیرد و با heap dump عیبش را پیدا کن.


بخش ۱۰ — چطور یک GC log را بخوانیم

JVMهای مدرن (۹+) از لاگِ یکپارچه (-Xlog:gc*) استفاده می‌کنند. یک توقفِ جوانِ G1 این شکلی است:

[2.335s][info][gc,start] GC(12) Pause Young (Normal) (G1 Evacuation Pause)
[2.340s][info][gc      ] GC(12) Pause Young (Normal) (G1 Evacuation Pause) 512M->48M(1024M) 5.219ms

این‌طور بخوانش: GC شمارهٔ ۱۲ یک توقفِ تخلیهٔ جوان بود؛ heap از ۵۱۲M قبل ← ۴۸M بعد رفت (یعنی ۴۶۴M آزاد شد)، کلِ heap ۱۰۲۴M است، و توقف ۵٫۲۱۹ms طول کشید.

دنبالِ این نشانه‌ها بگرد:

  • افزایشِ تعداد و طولِ توقف‌ها = دردسر در راه است.
  • دیدنِ Pause Full روی G1/ZGC = پرچمِ قرمز؛ بررسی کن (جهشِ تخصیص، آبجکت‌های humongous، heapِ کوچک).
  • بالارفتنِ پیوستهٔ اندازهٔ زنده پس از هر Full GC = نشتی (هر جمع‌آوری کمتر آزاد می‌کند؛ کف مدام بالا می‌رود).
  • to-space exhausted = آبجکت‌های جوان خیلی سریع ترفیع می‌یابند؛ اندازه‌ها را تیون کن.
ابزارها برای عیب‌یابیِ عمیق‌تر

برای دیدنِ محتوای heap یک dump بگیر (-XX:+HeapDumpOnOutOfMemoryError یا jmap -dump) و در Eclipse MAT بازش کن — «dominator tree» و «leak suspects»ِ آن دقیقاً می‌گویند چه چیزی حافظه را نگه داشته. برای تشخیصِ زنده: jstat -gcutil <pid> 1s، jcmd <pid> GC.heap_info و JFR (-XX:StartFlightRecording).


بخش ۱۱ — دام‌ها و نکاتِ ظریفِ رایج

این‌ها را زیاد اشتباه می‌کنند
  • صداکردنِ System.gc() — یک درخواست است نه دستور؛ معمولاً یک Full GCِ گران راه می‌اندازد و آسیب می‌زند. با -XX:+DisableExplicitGC خنثی‌اش کن.
  • finalize() — منسوخ، غیرقابلِ‌پیش‌بینی، می‌تواند آبجکت را دوباره زنده کند و جمع‌آوری را عقب بیندازد. از Cleaner یا try-with-resources استفاده کن.
  • object pooling برای آبجکت‌های کوچک — تخصیص در نسلِ جوان تقریباً رایگان است؛ pool‌کردن معمولاً با نگه‌داشتنِ آبجکت‌ها تا نسلِ قدیم فشارِ GC را بیشتر می‌کند. فقط منابعِ واقعاً گران (کانکشن، ترد) را pool کن.
  • -Xmxِ عظیم — heapِ بزرگ‌تر یعنی توقف‌های طولانی‌تر و Full GCهای دیرتر ولی بدتر. بزرگ‌تر همیشه بهتر نیست.
پرتگاهِ ۳۲GB (نکتهٔ سختِ سنیوری)

بالای ~۳۲GB، JVM دیگر compressed oops را نمی‌تواند استفاده کند، پس هر ارجاع از ۴ به ۸ بایت دوبرابر می‌شود. این سربار آن‌قدر زیاد است که یک heapِ ۳۱GB می‌تواند آبجکت‌های قابل‌استفادهٔ بیشتری از یک heapِ ۳۳GB نگه دارد! پس یا زیرِ ۳۲GB بمان، یا اگر رد می‌شوی، به‌قدرِ کافی رد شو که ارزشش را داشته باشد.


بخش ۱۲ — سؤالاتِ مصاحبه (با جواب)

۱) Collectorِ پیش‌فرض در Java 17 و 21 چیست؟

در هر دو G1. از Java 9 پیش‌فرض بوده (جایگزینِ Parallel شد). ZGC و Shenandoah اختیاری‌اند.

۲) آیا ZGC نسلی است؟ از کِی و چطور؟ (سنیور)

تا Java 20 تک‌نسلی بود. Generational ZGC در Java 21 آمد (JEP 439) ولی اختیاری با -XX:+UseZGC -XX:+ZGenerational. در Java 23 برای ZGC پیش‌فرض شد (JEP 474) و حالتِ غیرنسلی در Java 24 حذف شد (JEP 490).

۳) PermGen در برابر Metaspace — چه تغییر کرد و چرا مهم است؟

‏PermGen (پیش از Java 8) متادیتای کلاس را در ناحیه‌ای با اندازهٔ ثابت از heap نگه می‌داشت و باعثِ OutOfMemoryError: PermGen spaceِ مکرر در redeployها می‌شد. Java 8 آن را با Metaspace در حافظهٔ بومی که پویا رشد می‌کند جایگزین کرد. هنوز می‌شود نشتش داد (نشتیِ classloader)؛ با -XX:MaxMetaspaceSize محدودش کن.

۴) فرضیهٔ نسلی را توضیح بده و بگو چرا collectorهای کپی‌کننده کارآمدند.

بیشترِ آبجکت‌ها جوان می‌میرند. پس heap تقسیم می‌شود و young GC از یک collectorِ کپی‌کننده استفاده می‌کند که فقط آبجکت‌های زنده را لمس می‌کند (بازماندگان را کپی و Eden را یکجا reset می‌کند). چون بیشترِ آبجکت‌های جوان مرده‌اند، کارِ کمی می‌کند و رایگان فشرده‌سازی هم می‌شود — بدونِ fragmentation، با هزینه‌ای متناسبِ بازماندگان نه زباله.

۵) safepoint چیست و چرا حتی collectorهای «concurrent» هم STW دارند؟ (سنیور)

‏safepoint نقطه‌ای است که وضعیتِ ترد سازگار و برای JVM شناخته‌شده است تا بتوان بی‌خطر متوقفش کرد. حتی ZGC هم به توقف‌های کوتاهِ STW نیاز دارد (مثلاً شروع/پایانِ مارکِ ریشه‌ها) تا یک snapshotِ سازگار بسازد؛ اما بخشِ عمدهٔ مارک/جابه‌جایی concurrent است. «concurrent» یعنی بیشترِ کار با برنامه هم‌پوشانی دارد، نه صفر توقف.

۶) ZGC چطور مستقل از اندازهٔ heap به توقفِ زیرِ ۱ms می‌رسد؟ (سخت)

با اشاره‌گرهای رنگی (بیت‌های متادیتا داخلِ خودِ اشاره‌گر) به‌علاوهٔ load barrier: وقتی برنامه یک ارجاع را load می‌کند، barrier آن را چک/تصحیح می‌کند و به ZGC اجازه می‌دهد آبجکت‌ها را هم‌زمان با برنامه جابه‌جا کند. کارِ توقف متناسبِ تعدادِ ریشه‌های GC است (کران‌دار)، نه اندازهٔ heap — پس با رشدِ heap تا ترابایت، توقف‌ها مسطح می‌مانند.

۷) تحلیلِ فرار چیست و چه بهینه‌سازی‌هایی را ممکن می‌کند؟

‏C2 تحلیل می‌کند آیا آبجکت از متد/ترد فرار می‌کند. اگر نه: جایگزینیِ اسکالر (آبجکت روی heap ساخته نمی‌شود؛ فیلدها رجیستر/پشته می‌شوند ← صفر زباله) و حذفِ قفل (حذفِ synchronizedِ بی‌فایده). بهترین‌تلاش است، نه تضمین.

۸) tiered compilation و تقسیمِ C1/C2 را توضیح بده.

مفسر اول با پروفایل‌گیری اجرا می‌کند. C1 (سریع، بهینه‌سازیِ سبک) متدهای گرم را با پروفایلِ کامل کامپایل می‌کند (سطح ۳)؛ داغ‌ترین‌ها به C2 (کند، بهینه‌سازیِ تهاجمی → سطح ۴) می‌رسند. این استارتاپِ سریع (C1) را با اوجِ throughput (C2) متعادل می‌کند.

۹) نشتی را پیدا کن:
class Cache {
    private static final Map<Key, Value> M = new HashMap<>();
    static void put(Key k, Value v) { M.put(k, v); } // هرگز evict نمی‌شود
}

یک Mapِ static که فقط رشد می‌کند برای همیشه از یک لنگرِ GC قابل‌دسترس است ← رشدِ بی‌کرانِ heap ← OutOfMemoryError: Java heap space. اصلاح: محدودش کن (Caffeine/LRU) یا اگر باید خودکار جمع شوند از WeakHashMap/SoftReference استفاده کن.

۱۰) این چه چاپ می‌کند و دام کجاست؟ (نکتهٔ ظریف)
Integer a = 127, b = 127;
Integer c = 128, d = 128;
System.out.println((a == b) + " " + (c == d));

چاپ می‌کند: true false. متدِ Integer.valueOf بازهٔ ‎−۱۲۸ تا ۱۲۷ را کش می‌کند، پس a و b همان آبجکتِ کش‌شده‌اند (== درست است)، اما 128 خارجِ کش است ← دو آبجکتِ جدا (== نادرست). این autoboxing + کشِ Integer است و دقیقاً برای همین برای تایپ‌های باکس‌شده همیشه از .equals() استفاده می‌کنی.

۱۱) چرا گاهی دو کلاس با نامِ یکسان cast نمی‌شوند؟ (سنیور)

هویتِ زمان‌اجرای کلاس برابرِ (نام، ClassLoaderِ تعریف‌کننده) است. اگر دو ClassLoader هر کدام com.x.Foo را لود کنند، JVM آن‌ها را دو نوعِ متفاوت می‌بیند و castشان ClassCastException می‌دهد، حتی اگر سورس یکی باشد. رایج در app serverها، OSGi و hot-reload.

۱۲) سرویس در طولِ ساعت‌ها Full-GCِ مکرر و heapِ زندهٔ روبه‌بالا نشان می‌دهد. تشخیص؟ (سنیور)

امضای کلاسیکِ نشتیِ حافظه: هر Full GC کمتر آزاد می‌کند و کفِ زنده بالا می‌رود. heap dump بگیر، در Eclipse MAT باز کن، با dominator tree / leak suspects لنگرِ نگه‌دارنده را پیدا کن — اغلب یک کالکشنِ static، کشِ بی‌کران، یا ThreadLocalی که در pool مقدارش remove() نشده.

۱۳) چرا `-Xmx31g` می‌تواند از `-Xmx33g` بهتر باشد؟ (سخت)

بالای ~۳۲GB، JVM compressed oops را خاموش می‌کند، پس هر ارجاع از ۴ به ۸ بایت رشد می‌کند. این سربار می‌تواند ظرفیتِ قابل‌استفاده را چنان کم کند که heapِ ۳۱GB آبجکت‌های زندهٔ بیشتری از ۳۳GB نگه دارد. فقط وقتی واقعاً لازم است از مرز عبور کن.

۱۴) آیا `System.gc()` تضمین می‌کند جمع‌آوری اجرا شود؟

نه — یک اشاره است که JVM می‌تواند نادیده بگیرد، و -XX:+DisableExplicitGC آن را بی‌اثر می‌کند. وقتی هم اجرا شود معمولاً یک Full STW GCِ گران را اجبار می‌کند، پس تکیه بر آن در کدِ اپلیکیشن ضدالگوست.

۱۵) primitiveها، آبجکت‌ها و فیلدهای static کجا زندگی می‌کنند؟ (بنیادی اما اغلب اشتباه)

متغیرهای primitiveِ محلی و ارجاع‌های آبجکت روی پشتهٔ ترد (در فریم) هستند؛ خودِ آبجکت‌ها روی heap. متادیتای کلاس و فیلدهای static در Metaspace (بومی) هستند، ولی آبجکتی که فیلدِ static به آن اشاره می‌کند همچنان روی heap است. خلاصهٔ تیز: «primitiveها و ارجاع‌ها روی stack، آبجکت‌ها روی heap» — با این نکته که تحلیلِ فرار می‌تواند بعضی آبجکت‌های کوتاه‌عمر را کلاً خارج از heap نگه دارد.


نکاتِ سنیور و موارد پیشرفته

تا اینجا فهمیدیم حافظه کجا زندگی می‌کند و GC چطور تمیز می‌کند. اما در پروداکشن، چیزی که تو را ساعت سه صبح بیدار می‌کند معمولاً همین «لایهٔ پنهانِ» زیرِ آن مفاهیم است: چرا کانتینرت با اینکه -Xmx رعایت شده OOM-kill می‌شود، چرا یک ترد بی‌گناه کلِ JVM را قفل می‌کند، چرا کدی که «درست» است روی چند هسته نتیجهٔ کهنه می‌بیند. این بخش دقیقاً همان لایه است.

نقشهٔ راهِ این بخش

(۱) مدلِ حافظهٔ جاوا (JMM) — happens-before، volatile و انتشارِ امن؛ (۲) RSS در برابر heap و اینکه چرا کانتینر kill می‌شود؛ (۳) TLAB و مسیرِ واقعیِ تخصیص؛ (۴) card table و remembered set و ارجاع‌های بین‌نسلی؛ (۵) time-to-safepoint و دامِ حلقهٔ شمارشی؛ (۶) حالت‌های قفل در mark word؛ (۷) compressed class space، آبجکت‌های humongous، inline cache؛ (۸) startup و warmup (AppCDS/AOT/CRaC/GraalVM)؛ و ۸ سؤالِ سختِ سنیوری در انتها.


۱) مدلِ حافظهٔ جاوا (JMM): چیزی که فصل اصلی جا انداخت

فصل دربارهٔ «حافظه» بود، اما یک تکهٔ حیاتی‌اش کجاست؟ مدلِ حافظه. Heap بین همهٔ تردها مشترک است، ولی هر هسته cache و بافرِ نوشتنِ خودش را دارد و کامپایلر/CPU مجازند دستورها را جابه‌جا کنند. پس این سؤال اصلاً بدیهی نیست: «اگر ترد A فیلدی را بنویسد، ترد B کِی آن را می‌بیند؟»

دو کارمند و یک وایت‌بردِ مشترک

A روی دفترچهٔ شخصی‌اش (cacheِ هسته) می‌نویسد و فکر می‌کند B هم دیده. ولی تا وقتی روی وایت‌بردِ مشترک (حافظهٔ اصلی) کپی نکند، B نسخهٔ کهنه را می‌بیند. JMM قوانینِ «کِی باید روی وایت‌برد کپی کنی» است.

قلبِ JMM رابطهٔ happens-before است: اگر عملِ X، happens-before عملِ Y باشد، اثرِ X تضمیناً برای Y دیده می‌شود. مهم‌ترین یال‌ها:

  • قفلِ مانیتور: unlock روی یک قفل، happens-before هر lock بعدیِ همان قفل.
  • volatile: نوشتنِ یک فیلدِ volatile، happens-before هر خواندنِ بعدیِ همان فیلد.
  • شروع/پیوستنِ ترد: thread.start() happens-before کدِ داخلِ آن ترد؛ و کدِ ترد happens-before بازگشتِ thread.join().
  • تعدی (transitivity): اگر A→B و B→C پس A→C.
`volatile` سه چیز می‌دهد، یکی را نمی‌دهد

می‌دهد: دیده‌شدن (visibility)، جلوگیری از جابه‌جاییِ دستور (ordering)، و خواندن/نوشتنِ اتمیکِ long/double. نمی‌دهد: اتمیک‌بودنِ عملیاتِ مرکب. count++ روی فیلدِ volatile هم مسابقه‌ای است، چون read-modify-write است. برای شمارنده از AtomicInteger/LongAdder یا VarHandle (CAS) استفاده کن.

مسابقهٔ داده (data race) فقط «نتیجهٔ کهنه» نیست

دسترسیِ همزمانِ بدونِ همگام‌سازی که حداقل یک نوشتن دارد = data race و رفتارش تعریف‌نشده است؛ نه فقط ممکن است مقدار کهنه ببینی، بلکه ممکن است اصلاً هرگز آپدیت را نبینی (کامپایلر مقدار را در رجیستر hoist کند و حلقه‌ات ابدی شود). این باگ روی لپ‌تاپ x86 «کار می‌کند» و روی سرورِ ARM یا زیرِ بارِ سنگین می‌ترکد — بدترین نوعِ باگ.

انتشارِ امن (safe publication) و فیلدهای final. چطور یک آبجکت را بی‌قفل به تردِ دیگر بدهیم؟ JMM تضمین می‌کند اگر آبجکت درست ساخته شده باشد (یعنی this در حینِ سازنده فرار نکند)، فیلدهای final آن بلافاصله پس از پایانِ سازنده برای همهٔ تردها دیده می‌شوند — این «انجمادِ فیلدِ final» دقیقاً همان چیزی است که String و آبجکت‌های immutable را امن می‌کند: می‌توانی بی‌هیچ همگام‌سازی‌ای بینِ تردها پاسشان بدهی. راه‌های امنِ دیگر: نوشتن در فیلدِ volatile/AtomicReference، استفاده از static initializer (همان تضمینِ <clinit>)، یا عبور از یک collectionِ concurrent.

// double-checked locking درست: instance حتماً باید volatile باشد،
// وگرنه تردِ دیگر می‌تواند ارجاعِ غیرnull ولی نیمه‌ساخته ببیند.
class Lazy {
    private static volatile Lazy instance;   // بدونِ volatile شکسته است
    static Lazy get() {
        Lazy r = instance;
        if (r == null) synchronized (Lazy.class) {
            r = instance;
            if (r == null) instance = r = new Lazy();
        }
        return r;
    }
}
قضاوتِ سنیور

در عمل کمتر خودت DCL بنویس؛ الگوی Holder (که فصل نشان داد) ساده‌تر و بی‌خطاتر است. DCL را بدان تا در code review دام volatileِ فراموش‌شده را بگیری.


۲) RSS در برابر heap: چرا کانتینر با -Xmx2g در ۳GB kill می‌شود

شایع‌ترین «معمای پروداکشن»: heapِ زنده ۱.۲GB است، -Xmx2g گذاشته‌ای، اما کانتینر روی ۳GB به OOMKilled (exit 137) می‌خورد. چرا؟ چون heap فقط یک تکه از حافظهٔ کلِ پروسه است. kernel بر اساسِ RSS (حافظهٔ فیزیکیِ کلِ پروسه) می‌کُشد، نه بر اساسِ heapِ جاوا.

یک‌خط توضیح از تجزیهٔ فضای پروسه؛ RSS جمعِ همهٔ این‌هاست، نه فقط heap:

flowchart TD
  RSS["Container RSS  (what the kernel OOM-kills on)"] --> Heap["Java Heap (-Xmx)"]
  RSS --> Meta["Metaspace + Compressed Class Space"]
  RSS --> Code["JIT Code Cache"]
  RSS --> Stacks["Thread stacks  (N x -Xss)"]
  RSS --> Direct["Direct / mapped ByteBuffers (NIO, Netty)"]
  RSS --> GC["GC structures (card table, RSets, mark bitmaps)"]
  RSS --> Native["Native libs, JNI, malloc arenas"]
قاتلانِ خاموشِ حافظهٔ خارج از heap
  • Direct buffers / Netty: فریم‌ورک‌های شبکه آبجکت‌های DirectByteBuffer می‌سازند که خارج از heap‌اند و GC دیر آزادشان می‌کند؛ با -XX:MaxDirectMemorySize محدود کن.
  • پشتهٔ تردها: ۱۰۰۰ تردِ پلتفرمی × ۱MB = ۱GB حافظهٔ بومی که در -Xmx اصلاً دیده نمی‌شود.
  • glibc malloc arenas: روی لینوکس، تعدادِ زیادِ arena حافظهٔ RSS را باد می‌کند؛ MALLOC_ARENA_MAX=2 یک ترفندِ کلاسیکِ کاهشِ RSS در کانتینر است.
  • Metaspace: رشدِ بی‌سقف تا اتمامِ RAM.
ابزارِ درست: Native Memory Tracking

برای دیدنِ همهٔ تکه‌ها (نه فقط heap) با -XX:NativeMemoryTracking=summary استارت بزن و بعد jcmd <pid> VM.native_memory summary. این تنها راهِ اثباتِ اینکه «heap سالم است ولی Metaspace/Direct/Thread حافظه را می‌خورد» است. برای سایزینگِ درستِ کانتینر: -Xmx را روی حدودِ ۷۰–۷۵٪ حافظهٔ کانتینر بگذار (یا -XX:MaxRAMPercentage) و ~۲۵٪ را برای این حافظه‌های خارج از heap کنار بگذار.


۳) TLAB: تخصیص چطور «فقط جابه‌جاییِ یک اشاره‌گر» است

فصل گفت تخصیص در Eden «فقط bump کردنِ یک pointer» است. اما اگر همهٔ تردها روی یک pointerِ مشترک رقابت کنند، باید قفل بگیرند و کند می‌شود. راه‌حل: TLAB (Thread-Local Allocation Buffer). هر ترد یک تکهٔ اختصاصی از Eden می‌گیرد و داخلِ آن بدونِ هیچ همگام‌سازی‌ای فقط pointer را جلو می‌برد. برای همین تخصیصِ آبجکتِ کوچک در جاوا عملاً چند نانوثانیه است.

چرا این مهم است

وقتی آبجکت از TLAB بزرگ‌تر است یا TLAB پر شده، تخصیص به مسیرِ کند (قفلِ مشترک) یا مستقیم به old gen می‌افتد — «allocation outside TLAB». اگر پروفایلر (async-profiler حالتِ alloc) نرخِ بالای «allocation outside TLAB» نشان داد، یعنی آبجکت‌های خیلی بزرگ می‌سازی. این هم توضیح می‌دهد چرا escape analysis + scalar replacement انقدر قدرتمند است: آبجکتی که فرار نکند حتی وارد TLAB هم نمی‌شود.


۴) ارجاع‌های بین‌نسلی: card table و remembered set

اینجا یک تناقضِ ظاهری هست که سنیورها باید حلش را بلد باشند: young GC می‌خواهد فقط young را بگردد تا سریع باشد؛ اما اگر یک آبجکتِ old به یک آبجکتِ young اشاره کند، آن young زنده است — پس بدونِ اسکنِ کلِ old از کجا بفهمیم؟

card table و write barrier

راه‌حل: heap به «کارت‌های» ۵۱۲بایتی تقسیم می‌شود. هر بار که یک فیلدِ ارجاعی را می‌نویسی (a.f = b)، یک قطعه کدِ ریزِ مخفی به نامِ write barrier آن کارت را «کثیف» علامت می‌زند. حالا young GC فقط کارت‌های کثیف را به‌عنوانِ ریشهٔ اضافی می‌گردد، نه کلِ old را. G1 یک قدم جلوتر می‌رود: برای هر region یک remembered set (RSet) نگه می‌دارد که ثبت می‌کند کدام regionها به این region اشاره دارند — برای همین می‌تواند یک region را تنها جمع کند.

هزینهٔ پنهانِ G1

همان RSetها و write barrierها رایگان نیستند: هم CPU (روی هر نوشتنِ ارجاع) و هم حافظه مصرف می‌کنند. در برنامه‌هایی با نرخِ خیلی بالای موتیشنِ ارجاع (مثلاً گرافِ آبجکتِ بزرگ و پرتغییر)، سربارِ RSet یکی از دلایلی است که گاهی Parallel GC از G1 throughput بهتری می‌دهد. مارکِ همزمانِ G1 هم از یک write barrierِ دیگر به نامِ SATB استفاده می‌کند تا آبجکتی که در حینِ مارک ناپدید می‌شود را از دست ندهد.


۵) time-to-safepoint: چطور یک ترد کلِ JVM را قفل می‌کند

فصل گفت GC در safepoint همه را نگه می‌دارد. نکتهٔ سنیوریِ نادیده: رسیدنِ همه به safepoint فوری نیست. JVM «تعاونی» است؛ هر ترد فقط در نقاطِ خاصی (بازگشتِ متد، لبهٔ برگشتِ حلقه) safepoint را چک می‌کند. مشکل: HotSpot برای سرعت، در حلقه‌های شمارشیِ ساده (شمارندهٔ int با کرانِ مشخص) این چک را حذف می‌کند!

دامِ کلاسیکِ حلقهٔ شمارشیِ داغ

یک حلقهٔ سنگینِ for (int i=0; i<HUGE; i++) بدونِ فراخوانیِ متد ممکن است چند صد میلی‌ثانیه هیچ safepointی نزند. حالا اگر GC یا یک deopt بخواهد شروع شود، باید منتظرِ همهٔ تردها بماند — و این یک ترد کلِ JVM را نگه می‌دارد: pauseِ GC که باید ۵ms باشد، ناگهان ۳۰۰ms می‌شود. به این «Time-To-Safepoint spike» می‌گویند و در لاگ زیرِ [safepoint] (با -Xlog:safepoint) بخشِ reaching بالا می‌رود، نه بخشِ کارِ واقعیِ GC. علتِ دیگر: page faultِ سنگین یا JNI critical طولانی.

درمان

حلقه‌های شمارشیِ خیلی طولانی را بشکن، یا -XX:+UseCountedLoopSafepoints بده تا در لبهٔ برگشت هم poll بگذارد. اما اول اثبات کن که TTSP مشکل است (-Xlog:safepoint)، نه اینکه کورکورانه فلگ اضافه کنی.


۶) mark word فقط hashCode نیست: حالت‌های قفل

فصل mark word را «چیزهای مدیریتی» نامید. بازش کنیم، چون سؤالِ مصاحبه است. همان ۸ بایت، بسته به وضعیت، معنایِ متفاوت دارد و پایینش تگِ حالت است:

  • باز (unlocked): hashCode + بیت‌های سن (age).
  • قفلِ سبک (thin / lightweight): یک CAS اشاره‌گری به Lock Record روی پشتهٔ ترد می‌گذارد — قفلِ بی‌رقابتِ ارزان.
  • قفلِ سنگین (inflated / heavyweight): وقتی رقابت پیش می‌آید، قفل «باد می‌کند» و mark word به یک ObjectMonitor (mutex/park سطحِ OS) اشاره می‌کند.
  • علامتِ GC: در حینِ جمع‌آوری.
biased locking کجا رفت؟

سال‌ها یک حالتِ چهارم به نامِ biased locking بود (بهینه‌سازی برای قفلی که همیشه یک ترد می‌گیرد). در JDK 15 با JEP 374 به‌صورتِ پیش‌فرض غیرفعال و deprecated شد، چون پیچیدگی‌اش با الگوهای امروزی (thread poolها، lambdaها) نمی‌ارزید و «revokeِ bias» خودش TTSP spike می‌ساخت. سنیورِ به‌روز این را می‌داند: دیگر رویش حساب نکن.


۷) سه نکتهٔ ظریفِ دیگر که فصل باز نکرد

Compressed Class Space — «متاسپیسِ دوم». وقتی compressed class pointers روشن است (heap زیرِ ۳۲GB)، خودِ متادیتای klass در یک ناحیهٔ جدا و پیوسته به نامِ Compressed Class Space ذخیره می‌شود (پیش‌فرض ~۱GB رزرو، با -XX:CompressedClassSpaceSize). پس Metaspace در واقع دو تکه است و می‌توانی به‌طورِ خاص OutOfMemoryError: Compressed class space بگیری (جدا از Metaspace) — معمولاً وقتی هزاران کلاسِ داینامیک تولید می‌کنی.

آبجکت‌های humongous در G1

در G1 هر آبجکتی که از نصفِ اندازهٔ region بزرگ‌تر باشد «humongous» است، مستقیم در old gen و در regionهای پیوسته تخصیص می‌یابد. آرایه‌های بزرگِ اولیه (مثلاً byte[] چند مگابایتی) قاتلِ رایج‌اند: تخصیص و آزادسازی‌شان region را تکه‌تکه می‌کند و وقتی regionِ پیوستهٔ کافی پیدا نشود، Full GC راه می‌اندازد. اگر لاگت Pause Full و humongous allocation نشان داد، -XX:G1HeapRegionSize را بزرگ‌تر کن یا آبجکت‌های غول را بشکن.

Inline cache و چندریختیِ گران. JIT هر call siteِ مجازی را بر اساسِ تایپ‌هایی که دیده رنگ می‌کند: monomorphic (یک تایپ → inline + یک guard)، bimorphic (دو تایپ)، و megamorphic (بیش از دو → تسلیم می‌شود، vtable lookup می‌کند و inline نمی‌کند). یک interfaceِ خیلی پرمصرف با دهها پیاده‌سازی (مثلاً یک hot pathِ لاگینگ یا یک ابسترکشنِ عمومی) می‌تواند call siteها را megamorphic و کند کند — گاهی «کمتر ابسترکشن» در hot path واقعاً سریع‌تر است. مرتبط: OSR (On-Stack Replacement) اجازه می‌دهد یک حلقهٔ طولانی در حالِ اجرا کامپایل شود؛ و intrinsicها متدهایی مثل System.arraycopy، Math.max، Integer.bitCount هستند که JIT با اسمبلیِ دست‌نویس/دستورِ CPU جایگزین می‌کند.


۸) startup و warmup: هزینه‌ای که میکروسرویس‌ها را می‌سوزاند

JIT «گرم‌شدن» می‌خواهد: چند ثانیهٔ اول کد تفسیری و کند است. برای یک سرویسِ بلندمدت مهم نیست، اما برای serverless، scale-to-zero و اسکیلِ سریع در Kubernetes فاجعه است. راهکارهای مدرن:

  • AppCDS (Application Class Data Sharing): کلاس‌های پارس‌شده را در یک آرشیو ذخیره می‌کند تا استارتِ بعدی نپارسدشان.
  • JEP 483 (AOT Class Loading & Linking، JDK 24، پروژهٔ Leyden): یک قدم جلوتر از AppCDS؛ کلاس‌ها را از قبل load و link هم می‌کند و در «AOT cache» می‌گذارد؛ در دموی رسمی استارتِ Spring PetClinic تا ~۴۲٪ سریع‌تر شد. نیازمندِ یک «training run» شبیهِ پروداکشن است.
  • CRaC (Coordinated Restore at Checkpoint): از یک JVMِ گرم‌شده checkpoint می‌گیرد و در چند میلی‌ثانیه restore می‌کند — هم استارتِ فوری، هم peak performanceِ فوری (در بیلدهای Azul/Zulu موجود است).
  • GraalVM Native Image: کلِ برنامه را AOT به یک باینریِ بومی کامپایل می‌کند: استارتِ تقریباً آنی و حافظهٔ کم، اما بدونِ JIT (peakِ کمتر برای بارهای سنگین) و با دنیای بسته (reflection نیاز به کانفیگ دارد).
مصالحهٔ کلیدیِ JIT در برابر AOT

JIT در طولِ زمان با پروفایلِ واقعی به peak throughputِ بالا می‌رسد اما گرم‌شدن می‌خواهد. AOT (Native Image/Leyden) سریع استارت می‌زند اما ممکن است سقفِ throughput پایین‌تری داشته باشد. برای سرویسِ همیشه‌روشنِ پرترافیک → JIT/HotSpot؛ برای functionِ کوتاه‌عمر یا scale-to-zero → AOT/CRaC.

به‌روزرسانیِ مدرن: header فشرده حالا محصول است

فصل گفت compact object headers (JEP 450) در JDK 24 «آزمایشی» است. آپدیت: در JDK 25 با JEP 519 به یک قابلیتِ محصولِ کامل ارتقا یافت (mark word و klass pointer در یک واژهٔ ۶۴بیتیِ واحد ادغام می‌شوند؛ header از ۱۲ به ۸ بایت، تا ~۲۲٪ صرفه‌جوییِ heap در بنچمارک). هنوز پیش‌فرض نیست و با -XX:+UseCompactObjectHeaders روشن می‌شود، ولی فلگِ experimental حذف شده.


سؤالاتِ سختِ سنیوری (تکمیلی)

۱) `volatile` عملیاتِ `count++` را اتمیک می‌کند؟ اگر نه، چه چیزی سه‌گانهٔ visibility+ordering+atomicity را می‌دهد؟

نه. volatile فقط دیده‌شدن و ترتیب را تضمین می‌کند (و خواندن/نوشتنِ اتمیکِ long/double)، اما count++ یک read-modify-write است و دو تردِ همزمان می‌توانند آپدیتِ هم را گم کنند. برای اتمیک‌بودنِ خودِ عملیات به CAS نیاز داری: AtomicInteger/AtomicLong (یا زیرِ رقابتِ بالا LongAdder) یا VarHandle. جمله‌ای که باید بگویی: «volatile مشکلِ visibility را حل می‌کند، نه مشکلِ atomicity را.»

۲) چرا در double-checked locking فیلد باید `volatile` باشد؟

بدونِ volatile، تخصیصِ آبجکت سه مرحله است (تخصیصِ حافظه، اجرای سازنده، انتساب به فیلد) و JMM اجازهٔ جابه‌جایی این‌ها را می‌دهد. تردِ دوم می‌تواند ارجاعِ غیرnull را ببیند در حالی که سازنده هنوز کامل اجرا نشده — یعنی آبجکتِ نیمه‌ساخته. volatile این reorder را ممنوع و انتشار را امن می‌کند. (در عمل الگوی Holder را ترجیح بده که این دام را کلاً ندارد.)

۳) کانتینرت `-Xmx2g` دارد اما در ۳GB به OOMKilled (exit 137) می‌خورد، در حالی که هیچ `OutOfMemoryError`ای در لاگِ جاوا نیست. تشخیص؟

kernel بر اساسِ RSSِ کلِ پروسه می‌کشد، نه heapِ جاوا؛ و RSS = heap + Metaspace + code cache + پشتهٔ تردها + direct buffers + ساختارهای GC + malloc arenas. نبودِ OutOfMemoryError یعنی heap سالم است و مشکل خارج از heap است. با -XX:NativeMemoryTracking=summary + jcmd <pid> VM.native_memory تجزیه کن؛ مظنون‌ها: direct buffers (Netty)، نشتیِ ترد، Metaspace بی‌سقف، یا malloc arenaها (با MALLOC_ARENA_MAX=2 تست کن). درمان سایزینگ: -Xmx ~۷۵٪ حافظهٔ کانتینر و ~۲۵٪ برای بقیه.

۴) چطور یک ترد می‌تواند pauseِ GC را از ۵ms به ۳۰۰ms برساند بی‌آنکه GC کارِ بیشتری کند؟

از طریقِ time-to-safepoint. GC باید همهٔ تردها را در safepoint نگه دارد، اما HotSpot در حلقه‌های شمارشیِ سادهٔ int چکِ safepoint نمی‌گذارد. یک حلقهٔ داغِ طولانیِ بدونِ فراخوانیِ متد صدها میلی‌ثانیه به safepoint نمی‌رسد و همه منتظرش می‌مانند. در -Xlog:safepoint بخشِ «reaching» بالاست نه کارِ GC. درمان: شکستنِ حلقه یا -XX:+UseCountedLoopSafepoints — اما اول اثبات کن.

۵) young GC چطور بدونِ اسکنِ کلِ old gen می‌فهمد یک آبجکتِ young که فقط از old ارجاع دارد زنده است؟

با card table: heap به کارت‌های ۵۱۲بایتی تقسیم می‌شود و یک write barrier روی هر نوشتنِ فیلدِ ارجاعی کارتِ مربوطه را کثیف می‌کند. young GC فقط کارت‌های کثیفِ old را به‌عنوانِ ریشهٔ اضافی می‌گردد. G1 علاوه بر این برای هر region یک remembered set نگه می‌دارد. هزینه‌اش سربارِ write barrier و حافظهٔ RSet است — یکی از دلایلی که Parallel گاهی throughput بهتری از G1 دارد.

۶) biased locking چه بود و چه بر سرش آمد؟ حالت‌های قفل در mark word را بگو.

mark word حالتِ قفل را کد می‌کند: unlocked، thin/lightweight (CAS اشاره‌گر به Lock Record روی پشته، برای قفلِ بی‌رقابت)، و inflated/heavyweight (اشاره به ObjectMonitor سطحِ OS هنگامِ رقابت). biased locking حالتِ چهارمی بود برای قفلی که همیشه یک ترد می‌گرفت، اما در JDK 15 (JEP 374) پیش‌فرض غیرفعال و deprecated شد چون با thread poolها/lambdaهای امروزی نمی‌صرفید و «bias revocation» خودش safepoint/TTSP می‌ساخت.

۷) call siteِ megamorphic یعنی چه و چرا برای performance بد است؟

JIT هر call siteِ مجازی را بر اساسِ تعدادِ تایپ‌های دیده‌شده رنگ می‌کند: monomorphic (۱ تایپ → inline + guard)، bimorphic (۲)، megamorphic (بیش از ۲ → inline را رها می‌کند و vtable lookupِ گران می‌زند). یک interfaceِ خیلی مشترک با دهها پیاده‌سازی call siteها را megamorphic می‌کند و چون inlining از بین می‌رود، بهینه‌سازی‌های پس از inline هم از بین می‌روند. در hot path گاهی کاهشِ ابسترکشن (یا سیل‌کردنِ تایپ) واقعاً سریع‌تر است.

۸) میکروسرویس‌ات چند ثانیهٔ اولِ هر استارت کند است و در scale-to-zero درد می‌کشد. چه گزینه‌هایی داری و مصالحه‌شان چیست؟

علت: JIT گرم‌شدن می‌خواهد؛ کدِ اولیه تفسیری است. گزینه‌ها: AppCDS (آرشیوِ کلاس‌های پارس‌شده)؛ JEP 483 / AOT cache در JDK 24 که load+link را هم از قبل انجام می‌دهد (~۴۰٪ استارتِ سریع‌تر، نیازمندِ training run)؛ CRaC که از JVMِ گرم checkpoint/restore می‌کند (استارت و peakِ فوری)؛ و GraalVM Native Image که کلِ برنامه را AOT می‌کند (استارتِ آنی و حافظهٔ کم، اما بدونِ JIT پس peakِ کمتر و reflection نیازمندِ کانفیگ). قاعده: سرویسِ همیشه‌روشنِ پرترافیک → JIT؛ functionِ کوتاه‌عمر/scale-to-zero → AOT یا CRaC.

جمع‌بندیِ سنیورِ این بخش

heap تنها تکه از حافظهٔ پروسه است — کانتینر بر اساسِ RSS می‌کشد، پس NMT بلد باش. درستیِ همزمانی از JMM می‌آید: volatile = visibility+ordering (نه atomicity)، انتشارِ امن از فیلدِ final/volatile. تخصیص با TLAB سریع است؛ ارجاعِ بین‌نسلی با card table/RSet؛ و یک حلقهٔ شمارشیِ داغ از راهِ time-to-safepoint کلِ JVM را قفل می‌کند. biased locking رفته (JDK 15). و برای startup، دنیای مدرن از AppCDS → AOT (Leyden) → CRaC → Native Image انتخاب می‌کند و آگاهانه JIT-peak را با startup-speed تاخت می‌زند.

جمع‌بندیِ کلِ فصل

JVM یک کامپیوترِ خیالی است که بایت‌کدِ .class را اجرا می‌کند تا کدِ تو همه‌جا کار کند. ClassLoaderها کلاس‌ها را تنبل و در سه فاز (load → link → init) و با قانونِ «اول از والد بپرس» (برای امنیت) لود می‌کنند. حافظه در نواحیِ مشخص زندگی می‌کند: heap برای آبجکت‌ها، stack برای متغیرهای موقتِ هر ترد، metaspace برای متادیتای کلاس. Garbage Collector با پیدا‌کردنِ آبجکت‌های غیرقابل‌دسترس از لنگرها حافظه را خودکار آزاد می‌کند، و با تکیه بر «بیشترِ آبجکت‌ها جوان می‌میرند» heap را نسلی تقسیم می‌کند. G1 پیش‌فرضِ همه‌منظوره است؛ ZGC برای تأخیرِ فوق‌کم. JIT کدِ داغ را در لحظه به کدِ ماشین کامپایل می‌کند (C1 سریع، C2 تهاجمی). و بیشترِ باگ‌های حافظه در جاوا در واقع نشتی = قابلیتِ دسترسیِ ناخواسته هستند، نه کمبودِ واقعیِ رم.

منابع: JEP 439، JEP 474، JEP 490، Inside.java: Generational ZGC.

This chapter is the heart of Java. Get it right and the rest of the language becomes logical instead of memorized. So don't rush. We go step by step — first grab the idea with a simple analogy, then meet its technical name, then wire it back to real code and interview questions.

Roadmap for this chapter

First we understand what the JVM is and why it even exists. Then we open its three big parts: (1) how classes get loaded, (2) where memory lives (heap, stack, metaspace), (3) how the execution engine makes code fast (JIT) and keeps memory clean (Garbage Collector). We finish with tuning, memory leaks, reading a GC log, and interview questions.


Part 0 — Three words you must understand from scratch

Before anything, let's unpack three words that repeat throughout the chapter, in plain language. If these click, the rest is easy.

What does "abstraction" mean?

Imagine driving a car. You only deal with the steering wheel, gas pedal, and brake. What explosions happen inside the engine, how fuel burns, how the gears turn — none of that is your concern, and you don't need to know.

That's abstraction: "hide the complexity, expose only what the user needs."

What is a "virtual machine"?

An imaginary computer inside your real computer

The JVM is an imaginary, made-up computer running inside your real computer (Windows/Mac/Linux). You write your code for this "imaginary computer," not for Windows or Mac. The JVM abstracts (hides) the real hardware's complexity and translates your code into the real hardware's language.

The result: write your code once, run it everywhere — Windows, Mac, Linux, servers, phones. That's Java's famous slogan: Write once, run anywhere.

So when we say the JVM is an "abstract machine," we mean an imaginary computer that hides the underlying hardware details from you.

What does "stack-based" mean?

The JVM uses a structure called a stack to do its computations. A stack is exactly what it sounds like: a pile of plates. The last plate you put on is the first one you take off (this is called LIFO: Last In, First Out).

For example, to compute 2 + 3, the JVM does this: push 2 onto the stack, push 3 on top, then see the add instruction, pop both, add them, and push 5. You don't need to memorize the details right now; just know the stack is the JVM's moment-to-moment workbench.

The three key words

Abstraction = hiding complexity. Virtual machine = an imaginary computer that runs Java code so you don't depend on the hardware. Stack-based = a moment-to-moment workbench for computation, with "last in, first out" logic.


Part 1 — What exactly does the JVM do? (The big picture)

Let's trace the whole path from "code you write" to "what the CPU does" with an analogy.

The international restaurant

You write a recipe (your .java code). This recipe is in human English — readable to you, but hardware doesn't understand it.

A translator (called the compiler, javac) comes and translates your recipe into a standard international language (something like Esperanto). This intermediate language is called bytecode (the .class file).

Why do this? Because every chef understands this intermediate language — it doesn't matter if the chef is French (Windows) or Japanese (Linux). Each has a local translator called the JVM that takes the bytecode and runs it in its own hardware's language.

So the full path is:

your code (.java)  ──javac──▶  bytecode (.class)  ──JVM──▶  runs on real hardware
   readable English          standard intermediate language     the OS's native machine code

The subtle point: bytecode is platform-independent — the same .class file works on any system. The only thing that differs between operating systems is that last step (the JVM), which has a separate build for each OS.

The JVM does three things with bytecode:

  1. Load: find the .class file and bring it into memory.
  2. Verify: check that the bytecode is safe and well-formed.
  3. Execute: first interpret it line by line; and for the parts that repeat a lot ("hot code"), compile them to native machine code at runtime so they run fast. This runtime compilation is called JIT (Just-In-Time).

The JVM's three subsystems

Three departments of a big restaurant
  • Class Loader = procurement & storage: goes and finds the .class files, brings them in, and checks they're intact.
  • Runtime Data Areas = the physical kitchen space: the big fridge (heap), the moment-to-moment workbench (stack), the recipe library (metaspace) — where everything lives during work.
  • Execution Engine = the head chef and assistants: the one who actually cooks; includes the interpreter, the smart compiler (JIT), and the cleaner (Garbage Collector).

We'll open all three departments below.

The most important mindset shift: "standard" vs "product"

Here's a senior-level point many people get wrong:

Building code vs the contractor

The government writes a "building code": "every house must have a door, windows, a roof, and plumbing." The code doesn't say how to build it, only what output it must have. This is Java's Specification.

Now several contractors come and build houses following that code:

  • Oracle with a product called HotSpot (the most famous and the default) — packed with clever tricks for speed and memory management.
  • IBM with OpenJ9 — focused on lower RAM usage.
  • Azul with Zing — focused on very high scalability.

The conclusion here is interview gold:

Standard ≠ implementation

The JVM Specification = the code/standard (it defines behavior). HotSpot = Oracle's product that implements that standard. Almost everything we say in this chapter about GC, object headers, and JIT describes HotSpot's specific behavior, not a requirement of the spec. A junior thinks "Java = HotSpot." A senior knows Java is a standard and HotSpot is just one implementation of it — so when they hit a problem, they reach for HotSpot's specific tuning knobs rather than assuming it's inherent to the language.


Part 2 — Class loading: loading → linking → initialization

When your program runs, the JVM does not load all classes at once. It loads a class only when you actually need it for the first time. This is called lazy loading.

A 10-volume encyclopedia

When you buy a 10-volume encyclopedia, you don't read all 10 volumes right then! You only open the "letter K" volume when you actually need a word starting with K. The JVM is the same: it loads a class exactly when you first reach it (e.g., when you new it or call a static method). This makes the program start faster and avoids wasting memory.

When the JVM finally needs a class, it prepares it in three phases in a precise order (per JLS §12.4 and JVMS §5). Let's see them through the analogy of hiring a new employee.

Phase 1: Loading

The "HR department" (the ClassLoader) reads the bytes of the .class file (from disk, from inside a JAR, from the network, or even bytes generated on the fly) and builds a personnel file in memory — in Java this file is a Class<?> object placed on the heap.

A class's identity is not just its "name"

A class's runtime identity equals the combination of (fully-qualified name + the exact ClassLoader that loaded it). Meaning: if the exact same bytes are loaded by two different ClassLoaders, the JVM sees them as two completely different types!

The confusing error `X cannot be cast to X`

Because of that rule, you might see a strange error: ClassCastException: com.x.Foo cannot be cast to com.x.Foo — "Foo cannot be cast to Foo"! How? Because two different ClassLoaders loaded these two Foos, so to the JVM they're separate types. This trap is common in application servers, OSGi, and hot-reload.

Phase 2: Linking — three sub-steps

Now that the file is built, we connect the employee to the company's systems. Three sub-steps:

  1. Verification — the security check: the bytecode is checked for safety and correctness: is it type-safe? Are the stack operations valid? Does it have illegal jumps into the code? This is the backbone of JVM security — it's this verification that prevents tampered bytecode from corrupting memory.

  2. Preparation — the empty desk: static fields are created and given their default value (numbers 0, objects null, booleans false) — not their real value. Like giving the employee an empty desk and a blank notebook, not their actual tools.

  3. Resolution — turning promises into addresses: in your code you wrote "go work with the Database class." For now that's a symbolic name, vague. Here the JVM turns that name into a direct reference (the exact address of that thing in memory). This step can also be lazy and deferred to the last moment.

Phase 3: Initialization

Now it's "the first day of work." The JVM runs a special hidden method called <clinit> (short for class initializer). You don't write this method; the compiler builds it from two things and runs it top to bottom:

  • the real initialization of static fields (the 0s and nulls from the previous step now get their real values),
  • and the static { ... } blocks.

When does it run? Only on "active use" of the class: new, calling a static method, reading/writing a static field (unless it's a compile-time constant), reflection, or initializing a subclass.

The JVM's golden guarantee

The JVM guarantees <clinit> runs exactly once and in a fully thread-safe way. So if 100 threads reach this class for the first time simultaneously, the JVM queues them itself: only one initializes, the rest wait until it's ready. Without you writing a single synchronized keyword.

From this guarantee, a very popular and clever pattern for a lazy Singleton is built:

public class Config {
    private Config() {}                       // nobody outside can construct it

    // This inner class isn't loaded until getInstance() is called (lazy loading).
    private static class Holder {
        static final Config INSTANCE = new Config(); // <clinit> once, thread-safe, by the JVM
    }

    public static Config getInstance() { return Holder.INSTANCE; }
}
Why this pattern is genius

Until someone calls getInstance(), the Holder class doesn't even exist (lazy). The moment it's called, the JVM loads Holder, and because it guarantees <clinit> runs exactly once and safely, we've built a Singleton that is lazy, thread-safe, and lock-free — the best of both worlds.

ClassLoader hierarchy and the delegation model

The HR departments (ClassLoaders) work as a parent-child chain with one strict rule: "ask your boss first." This is the parent-delegation model: each loader first asks its parent to load the class; only if the parent fails does it try itself.

Bootstrap ClassLoader   (native/C++, loads the Java core: java.lang.* etc.)
   └─ Platform ClassLoader   (formerly "extension"; the JDK's own standard modules)
        └─ System/Application ClassLoader   (your program's classpath)
             └─ Custom / web-app ClassLoaders (e.g., Tomcat)

Why does this rule exist? Security. Let's see with an example:

The company's official seal

Suppose the CEO (Bootstrap) keeps the company's real, official seal in their safe — the world's most trusted seal. A hacker makes a fake seal that looks exactly like it and secretly places it on your desk (your project folder); they name it java.lang.String.

Without the delegation rule: you need a seal, you look at your own desk, grab the fake seal, and stamp the documents with it. The hacker wins — their virus-laced String code runs instead of Java's safe String, and the whole system is compromised.

With the delegation rule: you're not allowed to grab a seal from your desk on your own. You first ask the CEO (Bootstrap). They say, "I have the original String, use this." So the fake seal on your desk is never even seen. The hacker loses.

Why delegation brings security

Delegation prevents anyone from shadowing a core Java class (like java.lang.String) by dropping a same-named file in the project folder. Because the JVM always asks the core and upstream parents first, the original, safe version always wins.

Two version notes worth knowing

From Java 9 onward, the old "extension" loader became the Platform ClassLoader, and the whole thing is built on the module system (JPMS). Also, because Bootstrap is written in C++ and isn't a Java object, if you call SomeCoreClass.class.getClassLoader() you get null (meaning "I was loaded by Bootstrap").

Important exception: web servers like Tomcat invert the rule

In one Tomcat, 10 different web apps might run at once, each wanting a different version of a library (e.g., Log4j). If they all asked the parent, they'd collide. So Tomcat deliberately tells each web app's loader: "search your own folder first; if you don't find it, then ask the parent." This is the inverse of standard delegation, and its goal is to isolate apps so they don't break each other.


Part 3 — Runtime data areas: where does memory live?

This part shows you where each thing is stored when your program runs. Let's start with the kitchen analogy, then see the precise table.

The restaurant kitchen
  • Heap = the big fridge: everything you create with new (all objects and arrays) is kept here in bulk. This area is managed by the Garbage Collector.
  • Stack = the head chef's moment-to-moment workbench: each thread's small, temporary things (local variables, references, return address) go here and are cleared when the method finishes.
  • Metaspace = the recipe library: information about the classes themselves (not objects) lives here — structure, methods, bytecode.

Now the full table of areas and what error each throws when it fills up:

Area Shared among whom? What it holds Error on filling
Heap whole JVM all objects and arrays OutOfMemoryError: Java heap space
Metaspace whole JVM class metadata, method bytecode OutOfMemoryError: Metaspace
JVM Stack per thread frames (locals, operands, return address) StackOverflowError
PC Register per thread address of the current bytecode instruction
Native Method Stack per thread C/C++ code frames (via JNI)
Code Cache whole JVM JIT-compiled native code CodeCache is full (JIT turns off)

Two important notes about this table:

1. Each thread has its own stack. If a method calls itself infinitely (endless recursion), the stack fills up and you get StackOverflowError. On the other hand, if you create tens of thousands of platform threads, the system's native memory runs out (each stack is about 512KB–1MB, tunable with -Xss).

Why Virtual Threads (Java 21) matter

It's exactly this "each thread needs an expensive native stack" pressure that virtual threads solve: they don't permanently pin a native stack, so you can have millions of them. (They get their own chapter.)

2. Metaspace replaced PermGen. This is one of the most important Java 8 changes and a common interview question — we open it fully in the next section.

Zooming in on Metaspace: why and how it replaced PermGen

The car factory — the difference between Heap and Metaspace

Picture a car factory:

  • Heap = the big parking lot: every car built (say 10,000 of one model) is parked here. In Java, these cars are the objects.
  • Metaspace = the blueprint & engineering room: where the car's blueprint is kept.

The key point: to build 10,000 cars you need only one blueprint, not 10,000! In Java it's exactly the same — you make thousands of objects from one class (all in the Heap), but the structural info of the class itself is stored only once in Metaspace.

What exactly is in Metaspace? Metadata (i.e., "data about data"): the class name, the list of methods and their parameters, the list of fields and their types, inheritance info, and the constant pool. Note: the value of a field (e.g., that a user's name is "Ali") is on the Heap; but the fact that "the User class has a name field of type String" is in Metaspace.

What changed from PermGen to Metaspace?

Before Java 8 this space was called PermGen, lived inside the Heap, and had a fixed, small cap. If the program loaded many classes, it filled up and you got the infamous OutOfMemoryError: PermGen space. From Java 8, this space moved to the OS's native memory (off-heap) and was renamed Metaspace. The benefit: it now grows dynamically and doesn't crash on a small fixed cap until your physical RAM is full.

So why do we still see `OutOfMemoryError: Metaspace`?

Two main causes: (1) ClassLoader leak — especially in web servers: every time you redeploy code, Tomcat creates a new ClassLoader and loads all classes again; if the old ClassLoaders aren't discarded properly, duplicate copies pile up in Metaspace. (2) Dynamic class generation — libraries like Hibernate/Spring/CGLIB generate classes at runtime; if they run away, they fill Metaspace. That's why in production you cap it with -XX:MaxMetaspaceSize.


Part 4 — Object layout and headers (HotSpot, 64-bit)

Now that we know objects live on the Heap, let's see what an object exactly looks like in memory. Besides your own fields, each heap object has a header the JVM needs to manage it:

[ mark word: 8 bytes ][ klass pointer: 4 bytes ][ your fields... ][ padding to a multiple of 8 ]
  • Mark word — management stuff: identity hashCode, age bits for GC, lock state, and during a GC move, a forwarding pointer.
  • Klass pointer — points to the class metadata in Metaspace (i.e., "which class am I an instance of?").
What compressed oops is and why a pointer is 4 bytes

When the heap is smaller than ~32GB (the default), HotSpot uses compressed oops: it stores pointers in 4 bytes instead of 8 (as a scaled offset). This is a big memory saving — we come back to it in "why 31GB beats 33GB."

So an empty Object is 16 bytes (12 bytes header + 4 bytes padding). Arrays add an extra 4-byte "length" field.

Why this matters for performance

This header overhead explains why boxing is expensive: a boxed Integer takes about 16 bytes (header + value + padding) plus a separate reference, whereas a raw int is just 4 bytes. In hot loops and large data structures, this difference is really felt.

Looking ahead

Java 24's JEP 450 (compact object headers, still experimental) brings the header down to 8 bytes. Good to know, but not the default yet.


Part 5 — Garbage Collection fundamentals

Here we reach Java's magic: you don't free memory by hand; an automatic Garbage Collector does it. But how does it know which object is "garbage"?

GC works by "reachability," not counting

A string tied to an anchor

Imagine each object is a balloon and references are strings tying the balloons to each other and to a fixed anchor (a GC root). Roots are things that are always "alive": local variables of running threads, static fields, and JNI references.

The GC starts from the anchors and marks any balloon reachable through the strings as "alive." Any balloon with no path to an anchor is garbage and gets freed.

Why this beats "reference counting"

Because it handles cycles correctly. Suppose object A points to B and B points to A, but neither is tied to an anchor — an isolated "island." With naive reference counting, each one's counter is 1, so they're never freed (a leak!). But Java's tracing GC, starting from anchors and never reaching this island, correctly declares both dead.

The generational hypothesis

An empirical observation that the entire design of modern GCs is built on: most objects die young. That is, most objects become useless very soon after being created (like temporary variables inside a method). The GC exploits this fact by splitting the heap:

Young generation:  [ Eden | Survivor S0 | Survivor S1 ]      Old generation (Tenured)
  • New objects are created in Eden (ultra-fast allocation, just bumping a pointer).
  • A young GC (minor GC) copies the live objects of Eden and one Survivor into the other Survivor and increments their age. Since it only touches live objects and most are dead, it's very cheap.
  • An object that survives enough young cycles (-XX:MaxTenuringThreshold, up to 15) is promoted/tenured to the Old generation.
  • A major GC collects the old generation; and a Full GC collects everything (young + old + often metaspace). The Full GC is the expensive one you want to avoid.
What Stop-the-World means

So the GC can safely move objects and count anchors, the JVM briefly pauses all application threads at a safe point (safepoint) — this pause is called Stop-the-World (STW). Every GC has some STW. Modern GCs minimize it by doing most of the work concurrently with the program. Latency-sensitive systems live and die by the length and count of these pauses.

Weak references: weak / soft / phantom

Java has a few special reference types that let you tell the GC "this object isn't all that important":

  • SoftReference — cleared only under memory pressure. Good for caches (but be careful, it can leak).
  • WeakReference — cleared at the next GC once it's only weakly reachable. The basis of WeakHashMap.
  • PhantomReference — for deterministic cleanup of native resources after an object dies (the Cleaner class, the modern replacement for finalize()).

Part 6 — The GC landscape: which one, when?

Java has several different Garbage Collectors, and you pick one with a flag. Each is a trade-off between throughput (total work done) and latency (short pauses).

Collector Flag Pause model Best for Trade-off
Serial -XX:+UseSerialGC fully STW, single-threaded small heaps, single-CPU container, CLI tools simplest, lowest overhead; but long pauses at scale
Parallel -XX:+UseParallelGC fully STW, multi-threaded batch jobs, maximizing throughput best throughput; worst tail latency
G1 -XX:+UseG1GC (default since Java 9) mostly-concurrent mark, STW evacuation general purpose, large heap, latency/throughput balance you give a pause target (-XX:MaxGCPauseMillis, default 200ms)
ZGC -XX:+UseZGC concurrent, sub-1ms pauses very large heaps (up to terabytes), hard latency SLA slightly lower throughput, more CPU/memory overhead
Shenandoah -XX:+UseShenandoahGC concurrent, sub-1ms pauses low latency, Red Hat builds similar to ZGC
What's the default? (common interview question)

In all modern JDKs — including Java 17 and Java 21 — the default collector is still G1. ZGC and Shenandoah are opt-in. If someone asks "the default in 17 and 21?", the right answer: G1 (which replaced Parallel back in Java 9).

G1 with a bit of detail

G1 (short for Garbage-First) divides the heap into about 2048 equal regions (each 1–32MB). Each region dynamically gets tagged as Eden, Survivor, Old, or Humongous (objects larger than half a region). G1 marks mostly concurrently, then does STW evacuation pauses that pull live objects first out of the regions with the most garbage — hence "Garbage-First." You give a pause target, not an exact layout; G1 itself decides how many regions to collect to hit your target. G1 effectively eliminated the multi-second Full GCs of the old CMS collector (CMS was fully removed in Java 14).

ZGC and Generational ZGC — state the versions precisely

ZGC is a concurrent, region-based, compacting collector that, using tricks called colored pointers and load barriers, moves objects while the program is running. Its headline: pauses are sub-millisecond (typically 0.05–0.5ms) and do not grow with heap size — it scales to multi-terabyte heaps.

ZGC version history (memorize this precisely)
  • Java 15: ZGC became production-ready (JEP 377).
  • Java 21: added Generational ZGC (JEP 439), but opt-in: -XX:+UseZGC -XX:+ZGenerational. Plain -XX:+UseZGC gave the old non-generational version.
  • Java 23: generational mode became the default for ZGC (JEP 474), and the ZGenerational flag was deprecated.
  • Java 24: the non-generational mode was fully removed (JEP 490).

Stating these versions precisely shows the interviewer you actually follow the platform.

When ZGC instead of G1?

When the heap is large and you have a hard need for short tail latency (trading systems, ultra-low-latency services) where even G1's ~100–200ms pauses are too much. For general services, stay on G1 — it usually gives better throughput and less memory overhead, and it's the battle-tested default.


Part 7 — JIT compilation, tiered, and inlining

Remember we said the JVM first interprets code and compiles the hot parts? Let's make it precise.

HotSpot first interprets the bytecode and simultaneously profiles it: how many times was each method called? Which branches are taken most? Which types actually show up? Hot methods are turned into native code by two compilers:

  • C1 (client): fast compile, light optimization → fast startup.
  • C2 (server): slow compile, aggressive optimization (inlining, loop unrolling, dead-code elimination) → peak performance.
A rookie cook vs a genius head chef

The interpreter is like a rookie cook: it reads the bytecode line by line and executes it (normal speed, but it starts right now). The JIT is like a chef who notices which dish has been ordered a hundred times; they pre-convert that recipe into "automatic muscle memory" (native machine code) so next time it cooks at the speed of light.

Tiered compilation (the default) combines both compilers across 5 levels:

Level 0: interpreter
Level 1: C1, no profiling (trivial methods)
Level 2: C1, limited profiling
Level 3: C1, full profiling      ← most methods warm up here
Level 4: C2, fully optimized     ← the hottest methods reach here

Code starts interpreted, gets compiled by C1 after warming up with profiling, and the truly hot paths "graduate" to C2. Compiled code lives in the code cache.

A real production incident: the Code Cache filling up

If the code cache fills, the JIT turns off and everything falls back to the slow interpreter — a sudden performance cliff. The log says CodeCache is full. In very large programs you sometimes need to raise its cap (-XX:ReservedCodeCacheSize).

Inlining — the most valuable optimization

Inlining means replacing a call to a method with the body of that method. Why does it matter? Because once the body is copied in place, all the other optimizations can cross the method boundary and combine. C2 aggressively inlines small, hot methods.

Why "just add a getter" is free in Java

Many people fear adding getters/setters because they think of "method call cost." In practice C2 inlines these tiny methods and they vanish entirely — as if you'd touched the field directly. So don't worry about getter performance.

Escape analysis and scalar replacement

C2 tries to prove whether an object escapes the method/thread that created it (i.e., is seen anywhere outside). If it's proven never to escape:

  • Scalar replacement: the object is never allocated on the heap; its fields go straight into registers/the stack. That's why a hot loop creating a temporary Point can produce zero garbage.
  • Lock elision: if there's a synchronized on an object that doesn't escape, that lock is removed entirely.
// C2 can prove p doesn't escape; the temporary object may never reach the heap.
int sumOfSquares(int[] xs, int[] ys) {
    int total = 0;
    for (int i = 0; i < xs.length; i++) {
        Point p = new Point(xs[i], ys[i]); // scalar-replacement candidate
        total += p.x * p.x + p.y * p.y;
    }
    return total;
}
Don't rely on escape analysis

Escape analysis is not a guarantee — it's best-effort and can fail (e.g., when the object is passed to a non-inlined method). Don't design your program's correctness around it. And if you're benchmarking with JMH, know that this is exactly why JMH uses a Blackhole — to stop the compiler from deleting your "seemingly useless" code entirely.

Deoptimization

C2 makes speculative bets (e.g., "this call site has only ever seen ArrayList, so I'll assume it always does and generate optimized code"). If reality one day violates that assumption (suddenly a LinkedList arrives), the JVM throws away that compiled code, temporarily falls back to the interpreter, and may recompile. This shows up as a short performance dip after a new type appears.


Part 8 — Key flags every senior should know

# heap size — in production set min == max to avoid resize pauses and fragmentation
-Xms4g -Xmx4g

# container-aware (on by default since Java 10+): take heap as a percentage of container memory
-XX:MaxRAMPercentage=75.0

# collector choice
-XX:+UseG1GC            # default
-XX:+UseZGC             # generational from Java 23+
-XX:MaxGCPauseMillis=100

# per-thread stack size
-Xss512k

# Metaspace cap (otherwise it grows until native memory is exhausted)
-XX:MaxMetaspaceSize=256m

# on OOM: take a heap dump and die fast (for later diagnosis)
-XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/dumps
-XX:+ExitOnOutOfMemoryError

# GC log (unified, since Java 9+)
-Xlog:gc*:file=gc.log:time,uptime,level,tags:filecount=5,filesize=20m
Why `-Xms == -Xmx`?

If min and max differ, the committed heap keeps shrinking and growing, which itself causes Full GCs and page faults. Fixing the size removes that churn. In containers prefer MaxRAMPercentage so the JVM respects the cgroup limit rather than seeing the whole host's RAM and then getting OOM-killed.


Part 9 — Memory leaks and classifying OutOfMemoryError

What is a "memory leak" in Java?

In a GC language, a memory leak means unintended reachability: objects you're done with, but which are still reachable from a GC root, so the GC isn't allowed to free them and they pile up.

Classic leak sources:

  • static collections that only grow (static Map cache = ... that never removes anything).
  • Unbounded caches — use a size cap instead (Caffeine, LRU).
  • Listeners/callbacks that are never unregistered — the subject keeps the observer forever.
  • ThreadLocal in thread pools — a pooled thread never dies, so the ThreadLocal value is never freed; always remove() it in a finally.
  • ClassLoader leak — one lingering reference to a web app's classes keeps the whole ClassLoader (and all its classes in Metaspace) alive across redeploys.

Types of OutOfMemoryError (each means something different)

Message Meaning Common cause
Java heap space heap full, GC can't free enough leak or too-small heap
GC overhead limit exceeded >98% of time in GC, <2% freed heap nearly full, thrashing
Metaspace class metadata space exhausted classloader leak, too many dynamic classes
unable to create new native thread OS native memory or ulimit full thread leak, big -Xss × many threads
Direct buffer memory off-heap direct ByteBuffer exhausted leaking direct buffers
Requested array size exceeds VM limit requested an array near Integer.MAX_VALUE logic bug
Important: don't `catch` `OutOfMemoryError`

It's an Error, not an Exception. Catching it is almost always wrong, because the JVM may be in an unrecoverable state. Let the program die fast and find the bug with a heap dump.


Part 10 — How to read a GC log

Modern JVMs (9+) use unified logging (-Xlog:gc*). A G1 young pause looks like this:

[2.335s][info][gc,start] GC(12) Pause Young (Normal) (G1 Evacuation Pause)
[2.340s][info][gc      ] GC(12) Pause Young (Normal) (G1 Evacuation Pause) 512M->48M(1024M) 5.219ms

Read it like this: GC number 12 was a young evacuation pause; the heap went from 512M before → 48M after (so 464M was freed), the total heap is 1024M, and the pause took 5.219ms.

Look for these signs:

  • Rising count and length of pauses = trouble ahead.
  • Seeing Pause Full on G1/ZGC = a red flag; investigate (allocation spike, humongous objects, too-small heap).
  • The live size after each Full GC continually rising = a leak (each collection frees less; the floor keeps rising).
  • to-space exhausted = young objects are promoting too fast; tune the sizes.
Tools for deeper diagnosis

To inspect the heap's contents, take a dump (-XX:+HeapDumpOnOutOfMemoryError or jmap -dump) and open it in Eclipse MAT — its "dominator tree" and "leak suspects" tell you exactly what's holding memory. For live diagnosis: jstat -gcutil <pid> 1s, jcmd <pid> GC.heap_info, and JFR (-XX:StartFlightRecording).


Part 11 — Common pitfalls and subtleties

People get these wrong a lot
  • Calling System.gc() — it's a request, not a command; it usually triggers an expensive Full GC and hurts. Neutralize it with -XX:+DisableExplicitGC.
  • finalize() — deprecated, unpredictable, can resurrect an object and delay collection. Use Cleaner or try-with-resources.
  • Object pooling for small objects — allocation in the young generation is nearly free; pooling usually increases GC pressure by keeping objects alive into the old generation. Only pool genuinely expensive resources (connections, threads).
  • A huge -Xmx — a bigger heap means longer pauses and later-but-worse Full GCs. Bigger isn't always better.
The 32GB cliff (a hard senior nuance)

Above ~32GB, the JVM can no longer use compressed oops, so every reference doubles from 4 to 8 bytes. This overhead is so large that a 31GB heap can hold more usable objects than a 33GB heap! So either stay under 32GB, or if you cross it, cross it by enough to be worth it.


Part 12 — Interview questions (with answers)

1) What's the default collector in Java 17 and 21?

G1 in both. It's been the default since Java 9 (replaced Parallel). ZGC and Shenandoah are opt-in.

2) Is ZGC generational? Since when and how? (senior)

It was single-generation until Java 20. Generational ZGC arrived in Java 21 (JEP 439) but was opt-in with -XX:+UseZGC -XX:+ZGenerational. It became the default for ZGC in Java 23 (JEP 474), and non-generational mode was removed in Java 24 (JEP 490).

3) PermGen vs Metaspace — what changed and why does it matter?

PermGen (before Java 8) kept class metadata in a fixed-size region of the heap and caused frequent OutOfMemoryError: PermGen space on redeploys. Java 8 replaced it with Metaspace in native memory that grows dynamically. You can still leak it (classloader leak); cap it with -XX:MaxMetaspaceSize.

4) Explain the generational hypothesis and why copying collectors are efficient.

Most objects die young. So the heap is split, and the young GC uses a copying collector that only touches live objects (copies survivors and resets Eden in one shot). Since most young objects are dead, it does little work and compacts for free — no fragmentation, cost proportional to survivors not garbage.

5) What is a safepoint, and why do even "concurrent" collectors have STW? (senior)

A safepoint is a point where a thread's state is consistent and known to the JVM so it can be safely paused. Even ZGC needs short STW pauses (e.g., start/end of root marking) to build a consistent snapshot; but the bulk of marking/moving is concurrent. "Concurrent" means most of the work overlaps with the app, not zero pauses.

6) How does ZGC achieve sub-1ms pauses independent of heap size? (hard)

Via colored pointers (metadata bits inside the pointer itself) plus load barriers: when the app loads a reference, the barrier checks/fixes it, letting ZGC move objects concurrently with the app. The pause work is proportional to the number of GC roots (bounded), not the heap size — so as the heap grows to terabytes, pauses stay flat.

7) What is escape analysis and what optimizations does it enable?

C2 analyzes whether an object escapes the method/thread. If not: scalar replacement (the object isn't allocated on the heap; fields become registers/stack → zero garbage) and lock elision (removing useless synchronized). Best-effort, not a guarantee.

8) Explain tiered compilation and the C1/C2 split.

The interpreter runs first while profiling. C1 (fast, light optimization) compiles warm methods with full profiling (level 3); the hottest graduate to C2 (slow, aggressive optimization → level 4). This balances fast startup (C1) with peak throughput (C2).

9) Find the leak:
class Cache {
    private static final Map<Key, Value> M = new HashMap<>();
    static void put(Key k, Value v) { M.put(k, v); } // never evicts
}

A static Map that only grows is forever reachable from a GC root → unbounded heap growth → OutOfMemoryError: Java heap space. Fix: bound it (Caffeine/LRU), or use WeakHashMap/SoftReference if entries should be collectible.

10) What does this print, and where's the trap? (subtle)
Integer a = 127, b = 127;
Integer c = 128, d = 128;
System.out.println((a == b) + " " + (c == d));

Prints: true false. Integer.valueOf caches the range −128..127, so a and b are the same cached object (== true), but 128 is outside the cache → two distinct objects (== false). This is autoboxing + the Integer cache, and precisely why you always use .equals() for boxed types.

11) Why do two same-named classes sometimes fail to cast? (senior)

A class's runtime identity equals (name, defining ClassLoader). If two ClassLoaders each load com.x.Foo, the JVM sees them as two different types, and casting one to the other throws ClassCastException, even if the source is identical. Common in app servers, OSGi, and hot-reload.

12) A service shows frequent Full-GCs over hours and a live heap that keeps climbing. Diagnosis? (senior)

The classic signature of a memory leak: each Full GC frees less and the live floor rises. Take a heap dump, open it in Eclipse MAT, and use the dominator tree / leak suspects to find the retaining root — often a static collection, an unbounded cache, or a ThreadLocal not remove()d in a pool.

13) Why can `-Xmx31g` be better than `-Xmx33g`? (hard)

Above ~32GB, the JVM disables compressed oops, so every reference grows from 4 to 8 bytes. This overhead can reduce usable capacity so much that a 31GB heap holds more live objects than a 33GB one. Only cross the boundary when you truly need to.

14) Does `System.gc()` guarantee a collection runs?

No — it's a hint the JVM can ignore, and -XX:+DisableExplicitGC makes it a no-op. When honored it usually forces an expensive Full STW GC, so relying on it in application code is an anti-pattern.

15) Where do primitives, objects, and static fields live? (fundamental but often misstated)

Local primitive variables and object references live on the thread's stack (in the frame); the objects themselves on the heap. Class metadata and static fields are in Metaspace (native), while the object a static field points to is still on the heap. The crisp summary: "primitives and references on the stack, objects on the heap" — with the caveat that escape analysis can keep some short-lived objects off the heap entirely.


Senior notes & advanced edge cases

So far we know where memory lives and how GC cleans up. But in production, the thing that pages you at 3 a.m. is usually the hidden layer underneath those ideas: why your container gets OOM-killed even though -Xmx was respected, why one innocent thread freezes the whole JVM, why "correct" code sees stale values across cores. This section is exactly that layer.

Roadmap for this section

(1) The Java Memory Model — happens-before, volatile, safe publication; (2) RSS vs heap and why containers get killed; (3) TLAB and the real allocation path; (4) card tables & remembered sets for cross-generational references; (5) time-to-safepoint and the counted-loop trap; (6) lock states in the mark word; (7) compressed class space, humongous objects, inline caches; (8) startup & warmup (AppCDS/AOT/CRaC/GraalVM); and 8 hard senior interview questions at the end.


1) The Java Memory Model (JMM): the piece the main chapter skipped

The chapter was about "memory" — but where's its most critical piece, the memory model? The heap is shared across threads, yet each core has its own caches and write buffers, and the compiler/CPU are allowed to reorder instructions. So this question is not obvious at all: "if thread A writes a field, when does thread B see it?"

Two clerks and a shared whiteboard

A writes in his private notebook (the core's cache) and assumes B saw it. But until he copies it onto the shared whiteboard (main memory), B keeps reading the stale version. The JMM is the set of rules for when you must copy to the whiteboard.

The heart of the JMM is the happens-before relation: if action X happens-before action Y, then X's effects are guaranteed visible to Y. The edges that matter:

  • Monitor lock: an unlock on a lock happens-before every subsequent lock on the same lock.
  • volatile: a write to a volatile field happens-before every subsequent read of that field.
  • Thread start/join: thread.start() happens-before code inside that thread; that thread's code happens-before a returning thread.join().
  • Transitivity: if A→B and B→C then A→C.
`volatile` gives you three things — and denies one

Gives: visibility, ordering (no reordering across it), and atomic reads/writes of long/double. Denies: atomicity of compound operations. count++ on a volatile field is still racy, because it's a read-modify-write. For counters use AtomicInteger/LongAdder or a VarHandle (CAS).

A data race isn't just "a stale value"

Unsynchronized concurrent access where at least one is a write = a data race, and its behavior is undefined — you might not just read a stale value, you might never see the update at all (the compiler can hoist the field into a register and your loop spins forever). This bug "works" on your x86 laptop and explodes on an ARM server or under heavy load — the worst kind of bug.

Safe publication and final fields. How do you hand an object to another thread without a lock? The JMM guarantees that if an object is properly constructed (i.e., this does not escape during the constructor), its final fields are visible to all threads immediately after the constructor returns — this "final-field freeze" is exactly what makes String and immutable objects safe: you can pass them between threads with zero synchronization. Other safe ways: write through a volatile/AtomicReference field, use a static initializer (the <clinit> guarantee), or pass through a concurrent collection.

// Correct double-checked locking: instance MUST be volatile,
// otherwise another thread can see a non-null but half-constructed reference.
class Lazy {
    private static volatile Lazy instance;   // broken without volatile
    static Lazy get() {
        Lazy r = instance;
        if (r == null) synchronized (Lazy.class) {
            r = instance;
            if (r == null) instance = r = new Lazy();
        }
        return r;
    }
}
Senior judgment

In practice, rarely hand-write DCL — the Holder pattern (from the chapter) is simpler and unbreakable. Know DCL so you can catch the forgotten-volatile trap in code review.


2) RSS vs heap: why a -Xmx2g container gets killed at 3 GB

The most common "production mystery": live heap is 1.2 GB, you set -Xmx2g, yet the container hits OOMKilled (exit 137) at 3 GB. Why? Because the heap is only one slice of the whole process's memory. The kernel kills based on RSS (the process's total resident physical memory), not the Java heap.

A one-line decomposition; RSS is the sum of all of these, not just the heap:

flowchart TD
  RSS["Container RSS  (what the kernel OOM-kills on)"] --> Heap["Java Heap (-Xmx)"]
  RSS --> Meta["Metaspace + Compressed Class Space"]
  RSS --> Code["JIT Code Cache"]
  RSS --> Stacks["Thread stacks  (N x -Xss)"]
  RSS --> Direct["Direct / mapped ByteBuffers (NIO, Netty)"]
  RSS --> GC["GC structures (card table, RSets, mark bitmaps)"]
  RSS --> Native["Native libs, JNI, malloc arenas"]
The silent off-heap killers
  • Direct buffers / Netty: networking frameworks allocate DirectByteBuffers that live off-heap and are freed lazily by GC; cap them with -XX:MaxDirectMemorySize.
  • Thread stacks: 1000 platform threads x 1 MB = 1 GB of native memory that -Xmx never counts.
  • glibc malloc arenas: on Linux a large number of arenas bloats RSS; MALLOC_ARENA_MAX=2 is the classic container RSS-reduction trick.
  • Metaspace: uncapped, grows until RAM is gone.
The right tool: Native Memory Tracking

To see all the slices (not just heap), start with -XX:NativeMemoryTracking=summary, then jcmd <pid> VM.native_memory summary. This is the only way to prove "the heap is healthy but Metaspace/Direct/Thread memory is eating RAM." For correct container sizing: set -Xmx to ~70–75% of container memory (or -XX:MaxRAMPercentage) and leave ~25% for this off-heap memory.


3) TLAB: how allocation is "just bumping a pointer"

The chapter said Eden allocation is "just bumping a pointer." But if every thread contended on one shared pointer they'd need a lock and it would be slow. The fix: the TLAB (Thread-Local Allocation Buffer). Each thread gets its own private chunk of Eden and inside it just bumps a pointer with zero synchronization. That's why allocating a small object in Java is effectively a few nanoseconds.

Why this matters

When an object is bigger than the TLAB or the TLAB is full, allocation falls to the slow path (shared lock) or straight into the old gen — an "allocation outside TLAB." If a profiler (async-profiler in alloc mode) shows a high "allocation outside TLAB" rate, you're creating very large objects. This also explains why escape analysis + scalar replacement is so powerful: an object that never escapes never even touches a TLAB.


4) Cross-generational references: card tables and remembered sets

Here's an apparent contradiction seniors must be able to resolve: a young GC wants to scan only young to stay fast; but if an old object points to a young object, that young object is alive — so how do we know without scanning the whole old gen?

The card table and the write barrier

The fix: the heap is divided into 512-byte "cards." Every time you write a reference field (a.f = b), a tiny hidden snippet called the write barrier marks that card "dirty." Now the young GC only scans the dirty cards as extra roots, not the entire old gen. G1 goes one step further: it keeps a per-region remembered set (RSet) recording which regions point into this region — which is what lets it collect a single region in isolation.

G1's hidden cost

Those RSets and write barriers aren't free: they cost both CPU (on every reference write) and memory. In apps with a very high reference-mutation rate (a large, churny object graph), RSet overhead is one reason Parallel GC sometimes beats G1 on throughput. G1's concurrent marking also uses a different write barrier called SATB (snapshot-at-the-beginning) so it doesn't lose an object that disappears mid-mark.


5) Time-to-safepoint: how one thread freezes the whole JVM

The chapter said GC pauses everyone at a safepoint. The overlooked senior nuance: reaching a safepoint is not instant. The JVM is "cooperative"; each thread only checks for a safepoint at specific points (method returns, loop back-edges). The catch: for speed, HotSpot omits that check in simple counted loops (an int counter with a known bound)!

The classic hot counted-loop trap

A heavy for (int i=0; i<HUGE; i++) with no method call inside may go hundreds of milliseconds without hitting a single safepoint. Now if a GC or a deopt wants to start, it must wait for all threads — and this one thread stalls the whole JVM: a GC pause that should be 5 ms suddenly becomes 300 ms. This is a "time-to-safepoint spike," and in the log (-Xlog:safepoint) it shows up in the reaching portion, not in the actual GC work. Other causes: heavy page faults or a long JNI critical section.

The cure

Break up very long counted loops, or pass -XX:+UseCountedLoopSafepoints so a poll is placed on the back-edge too. But first prove TTSP is the problem (-Xlog:safepoint) rather than blindly adding flags.


6) The mark word isn't just hashCode: lock states

The chapter called the mark word "management stuff." Let's unpack it, because it's an interview question. Those same 8 bytes mean different things depending on state, with a state tag at the bottom:

  • Unlocked: hashCode + age bits.
  • Thin / lightweight lock: a CAS installs a pointer to a Lock Record on the thread's stack — a cheap uncontended lock.
  • Inflated / heavyweight lock: under contention the lock "inflates" and the mark word points to an OS-level ObjectMonitor (mutex/park).
  • GC-marked: during collection.
Where did biased locking go?

For years there was a fourth state, biased locking (an optimization for a lock always taken by one thread). In JDK 15, JEP 374 disabled it by default and deprecated it, because its complexity no longer paid off with modern patterns (thread pools, lambdas) and "bias revocation" itself caused TTSP spikes. An up-to-date senior knows: don't rely on it anymore.


7) Three more subtleties the chapter didn't open

Compressed Class Space — the "second Metaspace." When compressed class pointers are on (heap under 32 GB), the klass metadata itself is stored in a separate, contiguous region called the Compressed Class Space (default ~1 GB reserved, tuned with -XX:CompressedClassSpaceSize). So Metaspace is actually two parts, and you can specifically get OutOfMemoryError: Compressed class space (distinct from Metaspace) — usually when you generate thousands of dynamic classes.

Humongous objects in G1

In G1 any object larger than half a region is "humongous," allocated directly in the old gen across contiguous regions. Large primitive arrays (a multi-MB byte[]) are the usual culprits: allocating and freeing them fragments regions, and when enough contiguous regions can't be found, it triggers a Full GC. If your log shows Pause Full alongside humongous allocations, increase -XX:G1HeapRegionSize or break up the giant objects.

Inline caches and expensive polymorphism. The JIT colors each virtual call site by the types it has seen: monomorphic (one type → inline + a guard), bimorphic (two types), and megamorphic (more than two → it gives up, does a vtable lookup, and does not inline). A very heavily shared interface with dozens of implementations (a hot-path logging call, a generic abstraction) can turn call sites megamorphic and slow — sometimes "less abstraction" in a hot path really is faster. Related: OSR (On-Stack Replacement) lets a long-running loop be compiled while it's still executing; and intrinsics are methods like System.arraycopy, Math.max, Integer.bitCount that the JIT replaces with hand-written assembly/CPU instructions.


8) Startup & warmup: the cost that burns microservices

The JIT needs to "warm up": the first few seconds run interpreted and slow. For a long-lived service that's irrelevant, but for serverless, scale-to-zero, and fast Kubernetes scaling it hurts. The modern options:

  • AppCDS (Application Class Data Sharing): archives parsed classes so the next start doesn't re-parse them.
  • JEP 483 (AOT Class Loading & Linking, JDK 24, Project Leyden): one step past AppCDS — it also loads and links classes ahead of time into an "AOT cache"; the official demo shows Spring PetClinic starting up to ~42% faster. Requires a "training run" that mimics production.
  • CRaC (Coordinated Restore at Checkpoint): checkpoints an already-warmed JVM and restores it in milliseconds — instant startup and instant peak performance (available in Azul/Zulu builds).
  • GraalVM Native Image: compiles the whole program AOT into a native binary: near-instant startup and low memory, but no JIT (lower peak for heavy loads) and a closed world (reflection needs configuration).
The core JIT vs AOT trade-off

The JIT reaches high peak throughput over time using a real profile, but needs warmup. AOT (Native Image/Leyden) starts fast but may have a lower throughput ceiling. For an always-on, high-traffic service → JIT/HotSpot; for a short-lived function or scale-to-zero → AOT/CRaC.

Modern update: compact headers are now a product feature

The chapter said compact object headers (JEP 450) were "experimental" in JDK 24. Update: in JDK 25, JEP 519 promoted them to a full product feature (the mark word and klass pointer are merged into a single 64-bit word; the header drops from 12 to 8 bytes, up to ~22% heap savings in benchmarks). It's still not the default and is turned on with -XX:+UseCompactObjectHeaders, but the experimental flag is gone.


Hard senior interview questions (additional)

1) Does `volatile` make `count++` atomic? If not, what gives you visibility + ordering + atomicity all three?

No. volatile guarantees only visibility and ordering (plus atomic reads/writes of long/double), but count++ is a read-modify-write, so two concurrent threads can lose each other's update. For the operation itself to be atomic you need CAS: AtomicInteger/AtomicLong (or LongAdder under high contention) or a VarHandle. The line to say: "volatile solves visibility, not atomicity."

2) Why must the field be `volatile` in double-checked locking?

Without volatile, object construction is three steps (allocate memory, run the constructor, assign to the field) and the JMM permits reordering them. A second thread can see the non-null reference while the constructor hasn't finished — a half-constructed object. volatile forbids that reorder and makes the publication safe. (In practice prefer the Holder pattern, which doesn't have this trap at all.)

3) Your container has `-Xmx2g` but gets OOMKilled (exit 137) at 3 GB, and there's no `OutOfMemoryError` in the Java log. Diagnose it.

The kernel kills based on the whole process's RSS, not the Java heap; and RSS = heap + Metaspace + code cache + thread stacks + direct buffers + GC structures + malloc arenas. No OutOfMemoryError means the heap is healthy and the problem is off-heap. Break it down with -XX:NativeMemoryTracking=summary + jcmd <pid> VM.native_memory; suspects: direct buffers (Netty), a thread leak, uncapped Metaspace, or malloc arenas (test with MALLOC_ARENA_MAX=2). Sizing fix: -Xmx ~75% of container memory, ~25% for the rest.

4) How can one thread turn a GC pause from 5 ms into 300 ms without GC doing any more work?

Via time-to-safepoint. GC must pause all threads at a safepoint, but HotSpot places no safepoint check inside simple int counted loops. A long, hot loop with no method call goes hundreds of milliseconds without reaching a safepoint, and everyone waits for it. In -Xlog:safepoint the "reaching" portion is high, not the GC work. Cure: break up the loop or -XX:+UseCountedLoopSafepoints — but prove it first.

5) How does a young GC know a young object referenced only from old is alive, without scanning the whole old gen?

Via the card table: the heap is split into 512-byte cards and a write barrier dirties the relevant card on every reference-field write. The young GC scans only the dirty old-gen cards as extra roots. G1 additionally keeps a per-region remembered set. The cost is write-barrier overhead and RSet memory — one reason Parallel sometimes out-throughputs G1.

6) What was biased locking and what happened to it? Name the mark-word lock states.

The mark word encodes the lock state: unlocked, thin/lightweight (CAS a pointer to a stack Lock Record, for uncontended locks), and inflated/heavyweight (points to an OS-level ObjectMonitor under contention). Biased locking was a fourth state for a lock always taken by one thread, but in JDK 15 (JEP 374) it was disabled by default and deprecated because it no longer paid off with modern thread pools/lambdas and "bias revocation" itself caused safepoint/TTSP spikes.

7) What is a megamorphic call site and why is it bad for performance?

The JIT colors each virtual call site by how many types it has seen: monomorphic (1 type → inline + guard), bimorphic (2), megamorphic (more than 2 → it abandons inlining and does an expensive vtable lookup). A very heavily shared interface with dozens of implementations makes call sites megamorphic, and because inlining is lost, all the post-inline optimizations are lost too. In a hot path, reducing abstraction (or sealing the types) can genuinely be faster.

8) Your microservice is slow for the first few seconds of every start and it hurts under scale-to-zero. What options do you have and what's the trade-off?

The cause: the JIT needs to warm up; early code runs interpreted. Options: AppCDS (archive parsed classes); JEP 483 / AOT cache in JDK 24, which also load+links ahead of time (~40% faster startup, needs a training run); CRaC, which checkpoints/restores a warmed JVM (instant startup and peak); and GraalVM Native Image, which AOT-compiles the whole app (instant startup, low memory, but no JIT so lower peak and reflection needs config). Rule of thumb: always-on high-traffic service → JIT; short-lived/scale-to-zero function → AOT or CRaC.

The senior capsule for this section

The heap is only one slice of process memory — the container kills on RSS, so learn NMT. Concurrency correctness comes from the JMM: volatile = visibility+ordering (not atomicity), safe publication via final/volatile fields. Allocation is fast thanks to TLABs; cross-generational references are tracked by card tables/RSets; and a single hot counted loop can freeze the whole JVM via time-to-safepoint. Biased locking is gone (JDK 15). And for startup, the modern world chooses along AppCDS → AOT (Leyden) → CRaC → Native Image, deliberately trading JIT peak for startup speed.

The whole chapter in a nutshell

The JVM is an imaginary computer that runs .class bytecode so your code works everywhere. ClassLoaders load classes lazily, in three phases (load → link → init), following the "ask your parent first" rule (for security). Memory lives in defined areas: heap for objects, stack for each thread's temporary variables, metaspace for class metadata. The Garbage Collector frees memory automatically by finding objects unreachable from the roots, and splits the heap generationally by relying on "most objects die young." G1 is the general-purpose default; ZGC is for ultra-low latency. The JIT compiles hot code to machine code on the fly (C1 fast, C2 aggressive). And most memory bugs in Java are really leaks = unintended reachability, not a genuine shortage of RAM.

Sources: JEP 439, JEP 474, JEP 490, Inside.java: Generational ZGC.