Memory compaction function call tracing with ftrace and KGDB
-
Concept of Memory Compaction
![]() |
![]() |
|---|
- For Memory Stress, Using
stress-ng - https://github.com/ColinIanKing/stress-ng
sudo apt install stress-ng
-
Need to check
ftracefeature availability first -
Using
function_graphfor tracer -
TESTHistory-
`trace_memory_stress.sh`
- In the
trace_memory_stress.sh, it will trace thekcompactdduringstress-ng --vm 3 --vm-bytes 90% -t 10m & - You need to set the environment as follows first.
cd /sys/kernel/debug/tracing cat available_tracers echo function_graph > current_tracer ps -ef | grep kcompactd echo <PID> > /sys/kernel/debug/tracing/set_ftrace_pid
- Need to use 2 shells for shutting down the stress test process before 10minutes if you want or Change the minutes for execution.
- As we can see this image
(ftrace/trace_result.txt), there is no trace result. kcompactdis invoked mainly bykswapdand we test the program with allocating 90% of memory capacity- so probably
kswapdmust be executed and also it would invoke thekcompactd - but there is no result.....Hmm
- (Unlike kernel version with 5.4.0 which the proactive compaction would not be executed, 5.15 should be executed......)
- In the
-
`trace_manual_compaction.sh`
- If you did
trace_memory_stress.shtest, Need to make original state.echo > set_ftrace_pid
cat available_filter_functions echo "*compact*" > set_ftrace_filter echo "*migrate*" >> set_ftrace_filter
- Execute the
shell scriptandstress-ngseperately stress-ng --vm 1 --vm-bytes 90% -t 10mandsh trace_manual_compaction.sh
ftrace/trace_manual_OCI_ARM_result.txt
ftrace/trace_manual_swarm_result.txt
- In OCI, we can check that it tries to
migrate pagesfor memory compaction. - But in X86 Server, we can see
compact_unlock_should_abort.isra.0()but after this, we cannot see anymigrateor kind ofcompact_zonesymbol...... - See the reasons in KGDB below. In
compact_unlock_should_abortin ftrace test chapter
- If you did
-
- Comparing to 5.4.0(previous installed version in x86 server), lots of features are added.
proactive compaction, memory folios, ....
-
1. LAUNCH VM
launch vm
sudo apt-get install -y pkg-config libglib2.0-dev libpixman-1-dev libslirp-dev
# DOWNLOAD QEMU wget <https://download.qemu.org/qemu-8.1.2.tar.xz> tar xvJf qemu-8.1.2.tar.xz cd qemu-8.1.2 ./configure --enable-slirp make
# DOWNLOAD LINUX wget <https://cdn.kernel.org/pub/linux/kernel/v6.x/linux-6.6.tar.xz> tar -xvf linux-6.6.tar.xz #copy the config file cp linux-6.60.config linux-6.6/.config cd linux-6.6
make menuconfig # Load the config and Save and EXIT make -j$(nproc) # BUILD
#CREATE IMAGE with bootstrap cd .. chmod +x create_image.sh ./create_image.sh
./launch-vm.sh
- Actually, i tried to access the vm with
sshbut there seems error with network configuration. But i will pass this step (will FIX IT) ( THIS THING IS NOT THE PRIORITIZED for testing the compaction) - YOU CAN TURN OFF THE QEMU USING
Ctrl+a and x
- Actually, i tried to access the vm with
-
2. LAUNCH VM AND KGDB TOGETHER
LAUNCH VM AND KGDB TOGETHER
- It would be better to use TMUX
- You need at least 2 screen for KGDB in Host(Original) and for QEMU
- I just add one more screen to see the kernel source
- Procedure
- launch vm
- In host,
gdb linux-6.6/vmlinux- and
target remote [localhost:4321](http://localhost:4321)
- and
- you can use gdb commands.
-
compact_unlock_should_abortin ftrace testcompact_unlock_should_abort
- I debug
compact_unlock_should_abortfunction. ⇒ It was in ftrace test history above
/* * Compaction requires the taking of some coarse locks that are potentially * very heavily contended. The lock should be periodically unlocked to avoid * having disabled IRQs for a long time, even when there is nobody waiting on * the lock. It might also be that allowing the IRQs will result in * need_resched() becoming true. If scheduling is needed, compaction schedules. * Either compaction type will also abort if a fatal signal is pending. * In either case if the lock was locked, it is dropped and not regained. * * Returns true if compaction should abort due to fatal signal pending. * Returns false when compaction can continue. */ static bool compact_unlock_should_abort(spinlock_t *lock, unsigned long flags, bool *locked, struct compact_control *cc) { if (*locked) { spin_unlock_irqrestore(lock, flags); *locked = false; } if (fatal_signal_pending(current)) { cc->contended = true; return true; } cond_resched(); return false; } -FATAL SIGNAL Pending - Check if there is a pending signal in current process SIGKILL,...........
- FATAL SIGNAL Pending - Check if there is a pending signal in current process (SIGKILL,SIGTRAP…..)
static inline int task_sigpending(struct task_struct *p) { return unlikely(test_tsk_thread_flag(p,TIF_SIGPENDING)); } static inline int __fatal_signal_pending(struct task_struct *p) { return unlikely(sigismember(&p->pending.signal, SIGKILL)); } static inline int fatal_signal_pending(struct task_struct *p) { return task_sigpending(p) && __fatal_signal_pending(p); }
#0 compact_unlock_should_abort (cc=<optimized out>, locked=<optimized out>, flags=<optimized out>, lock=<optimized out>) at mm/compaction.c:569 #1 isolate_freepages_block (cc=cc@entry=0xffffc9000067fd00, start_pfn=start_pfn@entry=0xffffc9000067f958, end_pfn=end_pfn@entry=4456448, freelist=freelist@entry=0xffffc9000067fd00, stride=stride@entry=1, strict=strict@entry=false) at mm/compaction.c:614 #2 0xffffffff81325831 in isolate_freepages (cc=0xffffc9000067fd00) at mm/compaction.c:1711 #3 compaction_alloc (src=src@entry=0xffffea0004075e40, data=data@entry=18446683600576838912) at mm/compaction.c:1769 #4 0xffffffff813a2c0a in migrate_folio_unmap (ret=0xffffc9000067fad0, reason=MR_COMPACTION, mode=MIGRATE_ASYNC, dstp=<synthetic pointer>, src=0xffffea0004075e40, private=18446683600576838912, put_new_folio=0xffffffff81322ec0 <compaction_free>, get_new_folio=0xffffffff81325120 <compaction_alloc>) at mm/migrate.c:1123 #5 migrate_pages_batch (from=from@entry=0xffffc9000067fbb0, get_new_folio=get_new_folio@entry=0xffffffff81325120 <compaction_alloc>, put_new_folio=put_new_folio@entry=0xffffffff81322ec0 <compaction_free>, private=private@entry=18446683600576838912, mode=mode@entry=MIGRATE_ASYNC, reason=reason@entry=0, ret_folios=0xffffc9000067fad0, split_folios=0xffffc9000067fbd0, stats=0xffffc9000067fae4, nr_pass=3) at mm/migrate.c:1660 #6 0xffffffff813a3582 in migrate_pages_sync (from=from@entry=0xffffc9000067fbb0, get_new_folio=get_new_folio@entry=0xffffffff81325120 <compaction_alloc>, put_new_folio=put_new_folio@entry=0xffffffff81322ec0 <compaction_free>, private=private@entry=18446683600576838912, mode=mode@entry=MIGRATE_SYNC, reason=reason@entry=0, ret_folios=0xffffc9000067fbc0, split_folios=0xffffc9000067fbd0, stats=0xffffc9000067fbe4) at mm/migrate.c:1825 #7 0xffffffff813a4125 in migrate_pages (from=from@entry=0xffffc9000067fd10, get_new_folio=get_new_folio@entry=0xffffffff81325120 <compaction_alloc>, put_new_folio=put_new_folio@entry=0xffffffff81322ec0 <compaction_free>, private=private@entry=18446683600576838912, mode=MIGRATE_SYNC, reason=reason@entry=0, ret_succeeded=0xffffc9000067fcbc) at mm/migrate.c:1929 #8 0xffffffff81327b7a in compact_zone (cc=cc@entry=0xffffc9000067fd00, capc=capc@entry=0x0 <fixed_percpu_data>) at mm/compaction.c:2515 #9 0xffffffff81328536 in compact_node (nid=nid@entry=0) at mm/compaction.c:2812 #10 0xffffffff81328662 in compact_nodes () at mm/compaction.c:2825 #11 sysctl_compaction_handler (table=<optimized out>, buffer=<optimized out>, length=<optimized out>, ppos=<optimized out>, write=<optimized out>) at mm/compaction.c:2871 #12 sysctl_compaction_handler (table=<optimized out>, write=<optimized out>, buffer=<optimized out>, length=<optimized out>, ppos=<optimized out>) at mm/compaction.c:2858 #13 0xffffffff814a6bb7 in proc_sys_call_handler (iocb=<optimized out>, iter=0xffffc9000067fe58, write=write@entry=1) at fs/proc/proc_sysctl.c:600 #14 0xffffffff814a6cb3 in proc_sys_write (iocb=<optimized out>, iter=<optimized out>) at fs/proc/proc_sysctl.c:626 #15 0xffffffff813ef341 in call_write_iter (file=0xffff8881026d3200, iter=0xffffc9000067fe58, kio=0xffffc9000067fe80) at ./include/linux/fs.h:1956 #16 new_sync_write (ppos=0xffffc9000067fef0, len=2, buf=0x55555574aeb0 "1\n", filp=0xffff8881026d3200) at fs/read_write.c:491 #17 vfs_write (pos=0xffffc9000067fef0, count=2, buf=0x55555574aeb0 "1\n", file=0xffff8881026d3200) at fs/read_write.c:584 #18 vfs_write (file=0xffff8881026d3200, buf=0x55555574aeb0 "1\n", count=<optimized out>, pos=0xffffc9000067fef0) at fs/read_write.c:564 #19 0xffffffff813ef657 in ksys_write (fd=<optimized out>, buf=0x55555574aeb0 "1\n", count=2) at fs/read_write.c:637 #20 0xffffffff813ef70a in __do_sys_write (count=<optimized out>, buf=<optimized out>, fd=<optimized out>) at fs/read_write.c:649 #21 __se_sys_write (count=<optimized out>, buf=<optimized out>, fd=<optimized out>) at fs/read_write.c:646 #22 __x64_sys_write (regs=<optimized out>) at fs/read_write.c:646 #23 0xffffffff81e5193b in do_syscall_x64 (nr=<optimized out>, regs=0xffffc9000067ff58) at arch/x86/entry/common.c:50 #24 do_syscall_64 (regs=0xffffc9000067ff58, nr=<optimized out>) at arch/x86/entry/common.c:80 #25 0xffffffff820000d2 in entry_SYSCALL_64 () at arch/x86/entry/entry_64.S:120
compact_unlock_should_abortis called byisolate_freepages_block- When we see the image and the call stack, kind of
migrate_pagesorcompact*symbols should be detected.
- I debug
-
kcompactdfunctionkcompactd
#kcompactd function in Linux 6.6 #mm/compaction.c /* * The background compaction daemon, started as a kernel thread * from the init process. */ static int kcompactd(void *p) { pg_data_t *pgdat = (pg_data_t *)p; struct task_struct *tsk = current; long default_timeout = msecs_to_jiffies(HPAGE_FRAG_CHECK_INTERVAL_MSEC); long timeout = default_timeout; const struct cpumask *cpumask = cpumask_of_node(pgdat->node_id); if (!cpumask_empty(cpumask)) set_cpus_allowed_ptr(tsk, cpumask); set_freezable(); pgdat->kcompactd_max_order = 0; pgdat->kcompactd_highest_zoneidx = pgdat->nr_zones - 1; while (!kthread_should_stop()) { unsigned long pflags; /* * Avoid the unnecessary wakeup for proactive compaction * when it is disabled. */ if (!sysctl_compaction_proactiveness) timeout = MAX_SCHEDULE_TIMEOUT; trace_mm_compaction_kcompactd_sleep(pgdat->node_id); if (wait_event_freezable_timeout(pgdat->kcompactd_wait, kcompactd_work_requested(pgdat), timeout) && !pgdat->proactive_compact_trigger) { psi_memstall_enter(&pflags); kcompactd_do_work(pgdat); psi_memstall_leave(&pflags); /* * Reset the timeout value. The defer timeout from * proactive compaction is lost here but that is fine * as the condition of the zone changing substantionally * then carrying on with the previous defer interval is * not useful. */ timeout = default_timeout; continue; } /* * Start the proactive work with default timeout. Based * on the fragmentation score, this timeout is updated. */ timeout = default_timeout; if (should_proactive_compact_node(pgdat)) { unsigned int prev_score, score; prev_score = fragmentation_score_node(pgdat); proactive_compact_node(pgdat); score = fragmentation_score_node(pgdat); /* * Defer proactive compaction if the fragmentation * score did not go down i.e. no progress made. */ if (unlikely(score >= prev_score)) timeout = default_timeout << COMPACT_MAX_DEFER_SHIFT; } if (unlikely(pgdat->proactive_compact_trigger)) pgdat->proactive_compact_trigger = false; } return 0; }
-
Because of
Proactive Compaction,kcompactdshould be detected every 500ms. -
I add break point in
trace_mm_compaction_kcompactd_sleep(pgdat->node_id);line. Then it will break. (It is important to choose appropriate line for break point because it could be not detected.)
-
-
stress-ng+kcompactdfunction withprintkstress-ng
- I add lots of printk to check the compaction.
/* * The background compaction daemon, started as a kernel thread * from the init process. */ static int kcompactd(void *p) { pg_data_t *pgdat = (pg_data_t *)p; struct task_struct *tsk = current; long default_timeout = msecs_to_jiffies(HPAGE_FRAG_CHECK_INTERVAL_MSEC); long timeout = default_timeout; const struct cpumask *cpumask = cpumask_of_node(pgdat->node_id); if (!cpumask_empty(cpumask)) set_cpus_allowed_ptr(tsk, cpumask); set_freezable(); printk("1\n"); pgdat->kcompactd_max_order = 0; pgdat->kcompactd_highest_zoneidx = pgdat->nr_zones - 1; while (!kthread_should_stop()) { unsigned long pflags; printk("2\n"); /* * Avoid the unnecessary wakeup for proactive compaction * when it is disabled. */ if (!sysctl_compaction_proactiveness) timeout = MAX_SCHEDULE_TIMEOUT; printk("3\n"); trace_mm_compaction_kcompactd_sleep(pgdat->node_id); if (wait_event_freezable_timeout(pgdat->kcompactd_wait, kcompactd_work_requested(pgdat), timeout) && !pgdat->proactive_compact_trigger) { psi_memstall_enter(&pflags); kcompactd_do_work(pgdat); psi_memstall_leave(&pflags); /* * Reset the timeout value. The defer timeout from * proactive compaction is lost here but that is fine * as the condition of the zone changing substantionally * then carrying on with the previous defer interval is * not useful. */ timeout = default_timeout; printk("4\n"); continue; } /* * Start the proactive work with default timeout. Based * on the fragmentation score, this timeout is updated. */ timeout = default_timeout; if (should_proactive_compact_node(pgdat)) { unsigned int prev_score, score; printk("5\n"); prev_score = fragmentation_score_node(pgdat); proactive_compact_node(pgdat); score = fragmentation_score_node(pgdat); /* * Defer proactive compaction if the fragmentation * score did not go down i.e. no progress made. */ printk("6\n"); if (unlikely(score >= prev_score)) timeout = default_timeout << COMPACT_MAX_DEFER_SHIFT; } printk("7\n"); if (unlikely(pgdat->proactive_compact_trigger)) pgdat->proactive_compact_trigger = false; printk("8\n"); } printk("9\n"); return 0; }
stress-ng --vm 8 --vm-bytes 90% -t 10m- BUT IT ONLY prints 2,3,7,8,2,3,7,8,2,3,7,8…………
- Need to see
5,6for proactive compaction - I changed the vm or bytes several times.
- ALSO
cat /proc/vmstat
-
stress-ngcommands- Actually in Memory Compaction in Linux Kernel.pdf, the test scenario was done with
stress-ng - So i wanted to test it with same approach. (Of course, the kernel version is significantly different. 5.11 vs 6.6)
stress-ng analyze
-
Let’s check it from GDB.
gdb stress-ng run --vm 8 --vm-bytes 80% -t 10m
-
Because of
fork, it is detached. -
https://woosunbi.tistory.com/94 : Need to set child process debugging
-
BUT……………………..There is no symbol!I tried to compile the program with debug option(-g, -ggdb). But there are errors…….
-
FIX! (I modify the
stress-vecwide.c(took hours…..😱))stress-ng --vm 1 --vm-bytes 80% -t 10mstress_run_parallel→stress_run→rc = g_stressor_current->stressor->info->stressor(&args);:1439 → (stressor function)stress-vm.c : stress_vm()→stress_oomable_child→(func(args,context))stress_vm_child→stress_vm_all→mmap
buf = (uint8_t *)mmap(NULL, buf_sz, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS | vm_flags, -1, 0);
- It allocates memory with
mmap.- In
manpage andGNU documentation, flagMAP_ANONYMOUSis used forThis flag tells the system to create an anonymous mapping, not connected to a file. filedes and offset are ignored, and the region is initialized with zeros. Anonymous maps are used as the basic primitive to extend the heap on some systems. They are also useful to share data between multiple tasks without creating a file. On some systems using private anonymous mmaps is more efficient than using malloc for large blocks. This is not an issue with the GNU C Library, as the included malloc automatically uses mmap where appropriate. The mapping is not backed by any file; its contents are initialized to zero. The fd and offset arguments are ignored; however, some implementations require fd to be -1 if MAP_ANONYMOUS (or MAP_ANON) is specified, and portable applications should ensure this. The use of MAP_ANONYMOUS in conjunction with MAP_SHARED is supported on Linux only since kernel 2.4.
- In
-
- Let's check again about stress-ng.
- From the presentation pdf, it has too many features….
- Go to
manpage - Also, in example folder of stress-ng package , i can find some examples in
memory.jobandvm.job
#memory.job # # malloc stressor options: # start N workers continuously calling malloc(3), calloc(3), real‐ # loc(3) and free(3). By default, up to 65536 allocations can be # active at any point, but this can be altered with the --mal‐ # loc-max option. Allocation, reallocation and freeing are chosen # at random; 50% of the time memory is allocation (via malloc, # calloc or realloc) and 50% of the time allocations are free'd. # Allocation sizes are also random, with the maximum allocation # size controlled by the --malloc-bytes option, the default size # being 64K. The worker is re-started if it is killed by the out # of mememory (OOM) killer. # malloc 0 # 0 means 1 stressor per CPU # malloc-bytes 64K # maximum allocation chunk size # malloc-max 65536 # maximum number of allocations of chunks # malloc-ops 1000000 # stop after 1000000 bogo ops # malloc-thresh 1M # use mmap when allocation exceeds this size stress-ng --malloc 8 --malloc-bytes 1M --malloc-max 55000 -t 10m --metrics &
- kcompactd process works just for a moment
# # mmap stressor options: # start N workers continuously calling mmap(2)/munmap(2). The # initial mapping is a large chunk (size specified by # --mmap-bytes) followed by pseudo-random 4K unmappings, then # pseudo-random 4K mappings, and then linear 4K unmappings. Note # that this can cause systems to trip the kernel OOM killer on # Linux systems if not enough physical memory and swap is not # available. The MAP_POPULATE option is used to populate pages # into memory on systems that support this. By default, anonymous # mappings are used, however, the --mmap-file and --mmap-async # options allow one to perform file based mappings if desired. # mmap 0 # 0 means 1 stressor per CPU # mmap-ops 1000000 # stop after 1000000 bogo ops # mmap-async # msync on each page when using file mmaps # mmap-bytes 256M # allocate 256M per mmap stressor # mmap-file # enable file based memory mapping # mmap-mprotect # twiddle page protection settings
# # vm stressor options: # start N workers continuously calling mmap(2)/munmap(2) and writ‐ # ing to the allocated memory. Note that this can cause systems to # trip the kernel OOM killer on Linux systems if not enough physi‐ # cal memory and swap is not available. # vm 0 # 0 means 1 stressor per CPU # vm-ops 1000000 # stop after 1000000 bogo ops # vm-bytes 256M # size of each vm mmapping # vm-hang 0 # sleep 0 seconds before unmapping # vm-keep # keep mapping # vm-locked # lock pages into memory using MAP_LOCKED # vm-method all # vm data exercising method; use all types # vm-populate # populate (prefault) pages into memory stress-ng --vm 8 --vm-bytes 70% --vm-method all --vm-keep -t 10m score:10 stress-ng --vm 1 --vm-bytes 1G --vm-method all --vm-keep -t 10m score:30 stress-ng --vm 1 --vm-bytes 512M --vm-method all --vm-keep -t 10m score:39 stress-ng --vm 8 --vm-bytes 90% --vm-method all --vm-keep -t 10m score:0?????? stress-ng --vm 8 --vm-bytes 90% --vm-method all --vm-keep -t 10m score:0???? stress-ng --vm 8 --vm-bytes 90% -t 10m
#0 fill_contig_page_info (info=<synthetic pointer>, suitable_order=suitable_order@entry=9, zone=zone@entry=0xffff88843ffc8000) at mm/vmstat.c:1067 #1 extfrag_for_order (zone=zone@entry=0xffff88843ffc8000, order=order@entry=9) at mm/vmstat.c:1119 #2 0xffffffff813230fb in fragmentation_score_zone (zone=0xffff88843ffc8000) at mm/compaction.c:2100 #3 fragmentation_score_zone_weighted (zone=0xffff88843ffc8000) at mm/compaction.c:2117 #4 fragmentation_score_node (pgdat=pgdat@entry=0xffff88843ffc8000) at mm/compaction.c:2139 #5 0xffffffff81328b13 in should_proactive_compact_node (pgdat=0xffff88843ffc8000) at mm/compaction.c:2166 #6 kcompactd (p=0xffff88843ffc8000) at mm/compaction.c:3096
#fill_config_page_info /* │ * Calculate the number of free pages in a zone, how many contiguous │ * pages are free and how many are large enough to satisfy an allocation of │ * the target size. Note that this function makes no attempt to estimate │ * how many suitable free blocks there *might* be if MOVABLE pages were │ * migrated. Calculating that is possible, but expensive and can be │ * figured out from userspace │ */ #extfrag_for_order /* │ * Calculates external fragmentation within a zone wrt the given order. │ * It is defined as the percentage of pages found in blocks of size │ * less than 1 << order. It returns values in range [0, 100]. │ */ !!ORDER=9 #fragmentation_score_zone_weighted /* │ * A weighted zone's fragmentation score is the external fragmentation │ * wrt to the COMPACTION_HPAGE_ORDER scaled by the zone's size. It │ * returns a value in the range [0, 100]. │* │ * The scaling factor ensures that proactive compaction focuses on larger │ * zones like ZONE_NORMAL, rather than smaller, specialized zones like │ * ZONE_DMA32. For smaller zones, the score value remains close to zero, │ * and thus never exceeds the high threshold for proactive compaction. │ */ ZONE NORMAL에 신경써서 더한다. #fragmentation_score_node /* │ * The per-node proactive (background) compaction process is started by its │ * corresponding kcompactd thread when the node's fragmentation score │ * exceeds the high threshold. The compaction process remains active till │ * the node's score falls below the low threshold, or one of the back-off │ * conditions is met. │ */ *******score값이 >wmark_high 보다 클 때 작동 (리눅스 기본 default값이 90)
- Actually in Memory Compaction in Linux Kernel.pdf, the test scenario was done with
change
- It works... (it takes about 30s?)-
Now i can go into the condition
-
if(
should_proactive_compact_node) ->fragmentation_score_node-
In
fragmentation_score_node(), In general, it can be seen that it does not contribute to the score value unless it isZONE_NORMAL/* * A zone's fragmentation score is the external fragmentation wrt to the * COMPACTION_HPAGE_ORDER scaled by the zone's size. It returns a value * in the range [0, 100]. * * The scaling factor ensures that proactive compaction focuses on larger * zones like ZONE_NORMAL, rather than smaller, specialized zones like * ZONE_DMA32. For smaller zones, the score value remains close to zero, * and thus never exceeds the high threshold for proactive compaction. */ static unsigned int fragmentation_score_zone(struct zone *zone) { unsigned long score; score = zone->present_pages * extfrag_for_order(zone, COMPACTION_HPAGE_ORDER); return div64_ul(score, zone->zone_pgdat->node_present_pages + 1); } /* * The per-node proactive (background) compaction process is started by its * corresponding kcompactd thread when the node's fragmentation score * exceeds the high threshold. The compaction process remains active till * the node's score falls below the low threshold, or one of the back-off * conditions is met. */ static unsigned int fragmentation_score_node(pg_data_t *pgdat) { unsigned int score = 0; int zoneid; for (zoneid = 0; zoneid < MAX_NR_ZONES; zoneid++) { struct zone *zone; zone = &pgdat->node_zones[zoneid]; score += fragmentation_score_zone(zone); } return score; }
-
-
MEMORY ZONE: DMA32,DMA,NORMAL,MOVABLE,DEVICE (HIGHMEM for 32bit...)
- AFTER CHECKING
NORMAL, the score was 91.( > wmark_high)( ALMOST 90% of the score was from NORMAL)
- AFTER CHECKING
-
proactive_compact_node-
-->
copmact_zone -
compaction_suitable->isolate_miagratepages(isolate_migratepages_block) /migrate_pages$96 = {freepages = {next = 0xffffc9000026fde0, prev = 0xffffc9000026fde0}, migratepages = {next = 0xffffea0004001a48, prev = 0xffffea0004001a08}, nr_freepages = 0, nr_migratepages = 2, free_pfn = 4455936, migrate_pfn = 1049088, fast_start_pfn = 0, zone = 0xffff88843ffc8d00, total_migrate_scanned = 416, total_free_scanned = 0, fast_search_fail = 0, search_order = 0, gfp_mask = 3264, order = -1, migratetype = 0, alloc_flags = 0, highest_zoneidx = 0, mode = MIGRATE_SYNC_LIGHT, ignore_skip_hint = true, no_set_skip_hint = false, ignore_block_suitable = false, direct_compaction = false, proactive_compaction = true, whole_zone = true, contended = false, rescan = false, alloc_contig = false} ! number of migrate_pages 2 ! address would be 0xffffea0004001a08 - 0xffffea0004001a48 (sizeof(struct page)=0x40)
-
SUCCESS....
Unlike 5.11, there was noOOM
-
kswapdtokcompactdkswapd to kcompactd
-
As we saw in
ftraceresult, allocating almost 90% of memory would cause kswapd to reclaim pages and probably wake up kcompactd -
Command same as usual
-
And add breakpoint to
wakeup_kcompactd(mm/compaction.c:3017)
kswapdcallswake_up_kcompactd!- BUT the problem is that most of the allocating(reclaiming) order is 0, which means only 2^0=1 page so it has not woken up kcompactd.
- I waited for a long time just in case. But it was same.
-
Change the command using malloc.
-
stress-ng --malloc 8 --malloc-bytes 2M -t 10m
- order=3!! (2^3 pages)
- ->
kcompactd_node_suitable->compaction_suitable - The watermark was
COMPACT_SKIPPEDso we cannot see further procedure
-
- ftrace
- KERNEL+GDB
- Compaction & test & etc
- https://www.slideshare.net/AdrianHuang/memory-compaction-in-linux-kernelpdf
- https://hyeyoo.com/156
- https://elixir.bootlin.com/linux/v6.6/source/mm/compaction.c
- https://mitny.github.io/articles/2016-08/gdb-command
- https://mrrootable.tistory.com/28
- https://nitingupta.dev/post/proactive-compaction/
- https://medium.com/geekculture/linux-kernel-vs-memory-fragmentation-part-ii-9eab96030b6d













