-
Notifications
You must be signed in to change notification settings - Fork 69
Scheduling
CPU scheduling algorithms in MentOS.
The scheduler answers two separate questions, and it is worth keeping them apart from the very beginning:
-
Scheduling policy — which runnable task should run next? This is the algorithm
(Round-Robin, CFS, EDF, ...). In MentOS it lives in
kernel/src/process/scheduler_algorithm.c. -
Context switch — how do we stop one task and resume another? Saving and restoring
registers, switching the page directory, updating the TSS kernel stack. In MentOS this lives
in
kernel/src/process/scheduler.cand is the same code no matter which policy you pick.
Changing the policy changes nothing about the mechanism. Most of this page is about the policy.
MentOS lets you select a scheduling policy at configure time:
# kernel/CMakeLists.txt
set(SCHEDULER_TYPES SCHEDULER_RR SCHEDULER_PRIORITY SCHEDULER_CFS
SCHEDULER_EDF SCHEDULER_RM SCHEDULER_AEDF)
set(SCHEDULER_TYPE "SCHEDULER_RR" CACHE STRING "...")Selectable is not the same as implemented. In the current tree:
| Policy | CMake value | Status |
|---|---|---|
| Round-Robin |
SCHEDULER_RR (default) |
Complete. __scheduler_rr() is fully written. |
| Priority | SCHEDULER_PRIORITY |
Exercise scaffold. The comparison and the min initialiser are literal /*...*/ placeholders. |
| CFS | SCHEDULER_CFS |
Exercise scaffold. The vruntime comparison is a /*...*/ placeholder. |
| EDF | SCHEDULER_EDF |
Exercise scaffold. |
| RM | SCHEDULER_RM |
Exercise scaffold. |
| AEDF | SCHEDULER_AEDF |
Exercise scaffold. |
Each scaffold is written as:
static inline task_struct *__scheduler_cfs(runqueue_t *runqueue, bool_t skip_periodic)
{
#ifdef SCHEDULER_CFS
time_t min = /*...*/; // <- the part you are asked to write
/* ... */
return next;
#else
return __scheduler_rr(runqueue, skip_periodic);
#endif
}Selecting a scaffolded policy intentionally causes the build to fail at the exercise
placeholders until the missing implementation is completed. /*...*/ is a comment, so
time_t min = /*...*/; reaches the compiler as time_t min = ; — a syntax error. For example:
$ cmake .. -DSCHEDULER_TYPE=SCHEDULER_CFS && make
kernel/src/process/scheduler_algorithm.c:151:33: error: expected expression before ';' token
151 | time_t min = /*...*/;
| ^That error is the assignment, not a broken checkout. Throughout MentOS, /* ... */ marks a
deliberately incomplete section left for you to write; treat it as an exercise marker before
suspecting a defect.
Two details worth reading correctly in the snippet above:
- The
#else return __scheduler_rr(...)arm is not a runtime fallback. It is compiled only when that policy is not selected, and in that configuration the function is never called —scheduler_pick_next_task()dispatches through an#if defined / #elifchain. Its job is to keep the unselected functions well-formed. - Selecting no policy at all is a different, also-loud failure:
#error "You should enable a scheduling algorithm!".
This is deliberate: these are the exercises the course is built around. The algorithm descriptions below are therefore best read as specifications of what you are being asked to implement, with Round-Robin as the worked example.
Status: implemented and default.
__scheduler_rr().
Type: Time-sharing, preemptive
Best for: General-purpose, interactive systems
Algorithm:
- Each process gets a fixed time slice (quantum)
- Processes are arranged in a circular queue
- When quantum expires, process is preempted and moved to back of queue
- All processes get equal CPU time
Configuration:
cmake .. -DSCHEDULER_TYPE=SCHEDULER_RRCharacteristics:
- Fair: All processes get equal time
- Simple: Easy to implement and understand
- Low overhead: Minimal scheduling overhead
- Poor for real-time: No priority or deadline support
Time Quantum: Driven by timer ticks. MentOS programs the PIT with TICKS_PER_SECOND = 1193
(kernel/inc/hardware/timer.h), so a tick is roughly 0.84 ms — considerably finer than the
100 Hz / 10 ms figure textbooks usually quote.
Implementation: __scheduler_rr() in kernel/src/process/scheduler_algorithm.c. It walks
the runqueue starting from the node after the current task (wrapping once), and returns the
first TASK_RUNNING entry — which is exactly what makes it round-robin rather than
"always pick the head".
Status: exercise scaffold.
__scheduler_priority()has/*...*/placeholders; selectingSCHEDULER_PRIORITYmakes the build fail there until you complete it. The doc comment above it in the source spells out the pitfall you are meant to find.
Type: Priority-based, preemptive
Best for: Systems with varied workload priorities
Algorithm:
- Each process has a priority (lower value = higher priority)
- Scheduler always picks highest priority ready process
- Lower priority processes may starve if high-priority processes keep arriving
- Same priority processes use round-robin
Configuration:
cmake .. -DSCHEDULER_TYPE=SCHEDULER_PRIORITYCharacteristics:
- Prioritization: Important processes run first
- Responsive: High-priority tasks get CPU immediately
- Starvation risk: Low-priority processes may never run
- No fairness guarantees
Priority Levels:
-
0-99: Real-time priorities -
100-139: Normal priorities
Setting Priority:
#include <sched.h>
struct sched_param param;
param.sched_priority = 50;
sched_setparam(pid, ¶m);Status: exercise scaffold.
__scheduler_cfs()has/*...*/placeholders; selectingSCHEDULER_CFSmakes the build fail there until you complete it.
Type: Fair-share, time-based
Best for: Desktop/server systems, Linux-like behavior
Algorithm:
- Tracks "virtual runtime" (vruntime) for each process
- Always picks process with lowest vruntime
- vruntime increases as process runs
- Nice values affect vruntime increase rate
- Uses a runqueue scan to select the minimum vruntime task
Configuration:
cmake .. -DSCHEDULER_TYPE=SCHEDULER_CFSCharacteristics:
- Fair: Each process gets proportional CPU time
- Responsive: Good interactive performance
- Scalable: Efficient for many processes
- Complex: More sophisticated algorithm
Nice Values:
-
Range: -20 (highest priority) to +19 (lowest)
-
Default: 0
-
Setting nice value:
nice(10); // Decrease priority by 10
Status: exercise scaffold. Selecting
SCHEDULER_EDFmakes the build fail at the/*...*/placeholders in__scheduler_edf()until you complete it.
Type: Real-time, dynamic priority
Best for: Real-time systems with explicit deadlines
Algorithm:
- Each task has a deadline
- Scheduler always picks task with earliest deadline
- Optimal for single-CPU real-time scheduling
- Preempts if a task with earlier deadline arrives
Configuration:
cmake .. -DSCHEDULER_TYPE=SCHEDULER_EDFCharacteristics:
- Optimal: Maximizes number of tasks meeting deadlines
- Dynamic: Priorities change based on deadlines
- Real-time: Suitable for hard real-time systems
- Predictable: Behavior is deterministic
Setting Deadline:
struct sched_param param;
param.deadline = 1000; // absolute deadline in ticks
sched_setparam(pid, ¶m);Status: exercise scaffold. Selecting
SCHEDULER_RMmakes the build fail at the/*...*/placeholders in__scheduler_rm()until you complete it.
Type: Real-time, static priority
Best for: Periodic real-time tasks
Algorithm:
- Each task has a fixed period
- Priority is inversely proportional to period (shorter period = higher priority)
- Static priorities never change
- Preemptive
Configuration:
cmake .. -DSCHEDULER_TYPE=SCHEDULER_RMCharacteristics:
- Simple: Static priorities, easy to analyze
- Optimal: Among static-priority algorithms
- Predictable: Behavior is deterministic
- Limited: Only works for periodic tasks
Setting Period:
struct sched_param param;
param.period = 100; // period in ticks
sched_setparam(pid, ¶m);Status: exercise scaffold. Selecting
SCHEDULER_AEDFmakes the build fail at the/*...*/placeholders in__scheduler_aedf()until you complete it.
Type: Real-time, adaptive
Best for: Mixed periodic/aperiodic real-time workloads
Algorithm:
- Combines EDF with aperiodic task handling
- Periodic tasks use EDF
- Aperiodic tasks are scheduled in slack time
- Adapts to varying workload
Configuration:
cmake .. -DSCHEDULER_TYPE=SCHEDULER_AEDFCharacteristics:
- Flexible: Handles both periodic and aperiodic tasks
- Efficient: Maximizes CPU utilization
- Complex: More sophisticated than pure EDF
- Research-oriented: Experimental scheduler
| Scheduler | Type | Preemptive | Fair | Real-time | Complexity |
|---|---|---|---|---|---|
| RR | Time-sharing | Yes | Yes | No | Low |
| Priority | Priority | Yes | No | Partial | Low |
| CFS | Fair-share | Yes | Yes | No | Medium |
| EDF | Real-time | Yes | N/A | Yes | Medium |
| RM | Real-time | Yes | N/A | Yes | Low |
| AEDF | Real-time | Yes | Partial | Yes | High |
- Building a general-purpose system
- Want simplicity and fairness
- Interactive processes (shell, editor)
- Learning OS concepts
- Some tasks are more important than others
- Need explicit control over task importance
- Background tasks should yield to foreground
- Want Linux-like behavior
- Need fairness with some priority control
- Desktop or server workloads
- Many processes with varying importance
- Have real-time requirements with deadlines
- Tasks have explicit timing constraints
- Need optimal real-time scheduling
- Willing to specify deadlines explicitly
- All tasks are periodic
- Want static priority analysis
- Need predictable, deterministic behavior
- Simple real-time system
- Have both periodic and aperiodic real-time tasks
- Need flexible real-time scheduling
- Researching advanced scheduling algorithms
All schedulers implement:
// Initialize scheduler
void scheduler_initialize(void);
// Pick next task to run
task_struct *scheduler_pick_next_task(runqueue_t *runqueue);
// Enqueue a runnable task
void scheduler_enqueue_task(task_struct *task);
// Dequeue a task
void scheduler_dequeue_task(task_struct *task);
// Run scheduler (called from timer/interrupt paths)
void scheduler_run(pt_regs_t *f);- Source:
kernel/src/process/scheduler.c— runqueue management and context switching - Header:
kernel/inc/process/scheduler.h— also definesMAX_PROCESSES(256) - Algorithm implementations:
kernel/src/process/scheduler_algorithm.c— this is the file you edit when you implement a policy
// kernel/inc/process/scheduler.h
typedef struct runqueue {
size_t num_active; // number of queued processes
size_t num_periodic; // number of queued periodic processes
list_head_t queue; // the queue itself (intrusive list of task_struct)
task_struct *curr; // the currently running process
} runqueue_t;There is a single flat list, not one queue per priority level. A policy is therefore a
selection function over that list — which is why implementing a new policy means writing one
function, not restructuring the data. Tasks are linked in through task_struct::run_list, and
runqueue.curr is what __scheduler_rr() uses as its starting point.
num_periodic exists because MentOS supports periodic real-time tasks (see sched_setparam()
and waitperiod()); the skip_periodic argument threaded through the policy functions lets the
scheduler make a first pass that ignores them.
The scheduler is invoked from more places than the timer tick. scheduler_run() is called from:
-
Timer interrupt —
kernel/src/hardware/timer.c -
The
int 0x80handler —kernel/src/system/syscall.c, at the end of every system call, which is how a task that blocks in a syscall yields immediately -
The generic exception path —
kernel/src/descriptor_tables/exception.c -
The page-fault handler —
kernel/src/mem/page_fault.c -
Signal delivery —
kernel/src/system/signal.c
Saying "scheduling only happens on the timer tick" is therefore too strong.
Context switch process:
// Save current process state
scheduler_store_context(f, current_task);
// Pick next task
next_task = scheduler_pick_next_task(&runqueue);
// Restore next task state
scheduler_restore_context(next_task, f);
// Update current task pointer
current_task = next_task;MentOS bases scheduling decisions on timer ticks. TICKS_PER_SECOND is 1193
(kernel/inc/hardware/timer.h) — the PIT's base frequency is 1193182 Hz, and this is the divisor
MentOS picks.
Modify priority ranges:
// kernel/inc/process/prio.h
#define MAX_RT_PRIO 100
#define MAX_PRIO (MAX_RT_PRIO + NICE_WIDTH)// In kernel code
unsigned long start = timer_get_ticks();
scheduler_pick_next_task(&runqueue);
unsigned long end = timer_get_ticks();
pr_debug("Scheduler overhead: %lu ticks\n", end - start);These are indicative textbook figures for the algorithms, not measurements of the MentOS implementations (five of which are not implemented yet). Use them to reason about asymptotics, not to predict MentOS's numbers.
| Scheduler | Overhead (cycles) | O() complexity |
|---|---|---|
| RR | ~100 | O(1) |
| Priority | ~200 | O(1) |
| CFS | ~500 | O(log n) |
| EDF | ~400 | O(n) or O(log n) |
| RM | ~100 | O(1) |
| AEDF | ~600 | O(log n) |
// kernel/src/process/scheduler.c
#define __DEBUG_LEVEL__ LOGLEVEL_DEBUG
// Logs scheduler decisions
pr_debug("Switching from PID %d to PID %d\n", old_pid, new_pid);Use procfs entries such as /proc/<pid>/stat and scheduler feedback in /proc/feedback.
Test programs in userspace/tests/:
-
t_periodic1.c- Single periodic task -
t_periodic2.c- Multiple periodic tasks -
t_periodic3.c- Mixed periodic/aperiodic -
t_schedfb.c- Scheduler feedback test
Run tests:
make qemu
/bin/tests/t_periodic1MentOS keeps every policy in one file, so the cheapest way to add one is to follow the pattern the existing scaffolds use rather than creating a new source file:
-
Add a
__scheduler_custom()function inkernel/src/process/scheduler_algorithm.c, guarded the same way as the others:static inline task_struct *__scheduler_custom(runqueue_t *runqueue, bool_t skip_periodic) { #ifdef SCHEDULER_CUSTOM /* your selection logic; return a TASK_RUNNING entry from runqueue->queue */ #else return __scheduler_rr(runqueue, skip_periodic); #endif }
-
Add a branch for it in the
#if defined(SCHEDULER_RR) / #elif ...chain insidescheduler_pick_next_task(), in the same file. -
Append
SCHEDULER_CUSTOMtoSCHEDULER_TYPESinkernel/CMakeLists.txt. The existinglist(FIND ...)validation andtarget_compile_definitions()then handle the rest — you do not need to add a new source file to the build. -
Build:
cmake .. -DSCHEDULER_TYPE=SCHEDULER_CUSTOM make
Your function only has to answer "which task next?". Context saving/restoring, page-directory
switching and TSS updates are already handled by scheduler_run() and are policy-independent.
Goal: See which processes are scheduled and their CPU time.
Program - userspace/bin/sched_observer.c:
#include <stdio.h>
#include <unistd.h>
#include <sys/wait.h>
#include <time.h>
int main() {
printf("=== Scheduling Observer ===\n");
printf("Running 3 processes with same priority\n\n");
// Fork 3 child processes
for (int i = 0; i < 3; i++) {
pid_t pid = fork();
if (pid == 0) {
// Child: busy loop with periodic print
printf("[Child %d, PID %d] Started\n", i, getpid());
for (int j = 0; j < 5; j++) {
// Simulate work (tight loop)
volatile int sum = 0;
for (int k = 0; k < 100000; k++) {
sum += k;
}
printf("[Child %d, PID %d] Iteration %d\n", i, getpid(), j);
}
printf("[Child %d, PID %d] Done\n", i, getpid());
exit(0);
} else if (pid < 0) {
printf("Fork failed!\n");
return 1;
}
}
// Parent: wait for all children
printf("[Parent] Waiting for children...\n");
for (int i = 0; i < 3; i++) {
wait(NULL);
}
printf("[Parent] All children done\n");
return 0;
}Steps:
- Build:
make sched_observer - Run:
/bin/sched_observer -
Observe:
- Which child prints most often?
- How is CPU time distributed?
- Look for interleaved output (context switches)
Goal: Count how many times your process is rescheduled.
Program - userspace/bin/context_count.c:
#include <stdio.h>
#include <unistd.h>
#include <signal.h>
static volatile int switches = 0;
static volatile int iterations = 0;
void signal_handler(int sig) {
switches++;
// Signal indicates process was scheduled
}
int main() {
printf("=== Context Switch Counter ===\n");
printf("This process will count context switches\n");
printf("Use ps to see process switching\n\n");
// Install signal handler
signal(SIGALRM, signal_handler);
// Set alarm to fire every 100ms
alarm(1);
// Do work
for (int i = 0; i < 1000000; i++) {
iterations++;
// Simulate variable-length computation
for (volatile int j = 0; j < 100; j++);
}
printf("Completed %d iterations\n", iterations);
printf("Estimated context switches: %d\n", switches);
return 0;
}Advanced: Use /bin/ps to see process status during execution
Goal: Demonstrate how high-priority processes get more CPU time.
Program - userspace/bin/priority_test.c:
#include <stdio.h>
#include <unistd.h>
#include <sys/wait.h>
int main() {
printf("=== Priority Scheduling Test ===\n");
printf("If your kernel has priority scheduling,\n");
printf("high-priority processes should run more.\n\n");
// Create processes with different priorities
for (int priority = 0; priority < 3; priority++) {
pid_t pid = fork();
if (pid == 0) {
// Child process
// Try to set priority (if nice() is implemented)
if (priority > 0) {
nice(priority * 5); // Lower priority
}
printf("[PID %d, Priority %d] Starting work\n",
getpid(), priority);
// Do work
volatile long sum = 0;
for (int i = 0; i < 10000000; i++) {
sum += i;
}
printf("[PID %d, Priority %d] Finished\n",
getpid(), priority);
exit(0);
}
}
// Parent waits
for (int i = 0; i < 3; i++) {
wait(NULL);
}
printf("Test complete\n");
return 0;
}Observation:
- If priority scheduling works, process 0 (highest priority) finishes first
- Processes 1 and 2 may take much longer
Goal: Test different scheduling algorithms.
Steps:
-
Build with Round-Robin:
cd build cmake .. -DSCHEDULER_TYPE=SCHEDULER_RR make make qemu -
Run your own test program and observe timing (there is no
/bin/sched_observerin the tree — use the program from Exercise 1, or one of thet_periodic*tests) -
Rebuild with Priority scheduler:
cmake .. -DSCHEDULER_TYPE=SCHEDULER_PRIORITY make make qemu
-
Run same test, compare behavior
Analysis:
- RR: More balanced CPU distribution
- Priority: High-priority processes get more CPU
What happens today: step 3 does not get as far as running. __scheduler_priority() is still
an exercise scaffold, so configuring with SCHEDULER_PRIORITY makes the build stop at its
/*...*/ placeholders:
kernel/src/process/scheduler_algorithm.c:107:33: error: expected expression before ';' tokenImplementing __scheduler_priority() so that this exercise can be completed — and then produces
the difference described above — is the assignment.
Goal: See how preemption works in real-time.
Program - userspace/bin/preemption_demo.c:
#include <stdio.h>
#include <unistd.h>
#include <sys/wait.h>
#include <signal.h>
static volatile int preempted = 0;
void preempt_handler(int sig) {
preempted++;
printf(" [Preempted! Count: %d]\n", preempted);
}
int main() {
printf("=== Preemption Demo ===\n");
printf("This shows how kernel preempts long-running processes\n\n");
// Install signal to catch preemption
signal(SIGALRM, preempt_handler);
alarm(1); // Signal every second
printf("[PID %d] Starting long computation...\n", getpid());
printf("Watch how I'm interrupted!\n\n");
// Very long loop with prints
for (int i = 0; i < 50; i++) {
printf("Iteration %d\n", i);
// Simulate work
for (volatile int j = 0; j < 50000000; j++);
sleep(0); // Yield to other processes
}
printf("\nDone! Was preempted %d times\n", preempted);
return 0;
}Advanced Goal: Understand how the scheduler maintains fairness
Research:
- Read
kernel/src/process/scheduler.cto understand current algorithm - Look at process queue management
- Challenge: Modify scheduler to prioritize I/O-bound processes
- Test with programs that mix computation and I/O
- Linux Scheduler Documentation
- Real-Time Scheduling
- Process Management - Process lifecycle
- Development Guide - Implementing features
Previous: Debugging | Next: Contributing →