[LTP] [PATCH v2] getrusage03: account for percpu RSS counter batching

Nirmoy Das nirmoyd@nvidia.com
Wed Sep 16 17:58:41 CEST 2026


Hi Cyril,

I tested Jan’s patch on a machine  with 352 online CPUs and 64 KiB pages. Pinning was confirmed, but both the
baseline and Jan’s patch failed 10/10 runs:

child.children = 270336, expected 307200

Pinning keeps the allocation updates on one CPU, but get_mm_counter()
still reads only the global count. One batch here is
max(32, 352 * 2) = 704 pages, or 45056 KiB. The 36864 KiB shortfall
fits within that batch but exceeds the existing 20480 KiB DELTA_MAX.

I also tested a small follow-up that increases only the lower allowance
for the 100, 300 and 400 MiB checks. It passed 10/10 runs, with all nine
subtests passing each time.

I will send this follow-up once Jan’s patch is merged.

Regards,
Nirmoy

From: Cyril Hrubis <chrubis@suse.cz>
Date: Tuesday, 15. September 2026 at 17:29
To: Nirmoy Das <nirmoyd@nvidia.com>
Cc: ltp@lists.linux.it <ltp@lists.linux.it>; Jan Stancek <jstancek@redhat.com>
Subject: Re: [PATCH v2] getrusage03: account for percpu RSS counter batching

[You don't often get email from chrubis@suse.cz. Learn why this is important at https://aka.ms/LearnAboutSenderIdentification ]

Hi!
> diff --git a/testcases/kernel/syscalls/getrusage/getrusage03.c b/testcases/kernel/syscalls/getrusage/getrusage03.c
> index a2cdd6158..38a100576 100644
> --- a/testcases/kernel/syscalls/getrusage/getrusage03.c
> +++ b/testcases/kernel/syscalls/getrusage/getrusage03.c
> @@ -13,9 +13,13 @@
>   * this program.
>   */
>
> +#define _GNU_SOURCE
>  #include <stdlib.h>
>  #include <stdio.h>
>
> +#include "lapi/cpuset.h"
> +#include "tst_safe_stdio.h"
> +#include "tst_cpu.h"
>  #include "tst_test.h"
>  #include "getrusage03.h"
>
> @@ -23,12 +27,99 @@
>
>  static struct rusage ru;
>  static long maxrss_init;
> +static long lower_allowance;
>
>  static const char *const resource[] = {
>       TESTBIN,
>       NULL,
>  };
>
> +static long count_online_cpus(void)
> +{
> +     FILE *fp = SAFE_FOPEN("/proc/stat", "r");
> +     char line[BUFSIZ];
> +     long count = 0;
> +
> +     while (fgets(line, sizeof(line), fp)) {
> +             if (line[0] == 'c' && line[1] == 'p' && line[2] == 'u' &&
> +                 line[3] >= '0' && line[3] <= '9')
> +                     count++;
> +     }
> +
> +     if (ferror(fp))
> +             tst_brk(TBROK | TERRNO, "fgets(/proc/stat)");
> +
> +     SAFE_FCLOSE(fp);
> +
> +     if (!count)
> +             tst_brk(TBROK, "No online CPUs found in /proc/stat");
> +
> +     return count;
> +}
> +
> +static void pin_to_cpu(void)
> +{
> +     long ncpus = tst_ncpus_max();
> +     size_t size = CPU_ALLOC_SIZE(ncpus);
> +     cpu_set_t *mask = CPU_ALLOC(ncpus);
> +     int cpu = -1;
> +
> +     if (!mask)
> +             tst_brk(TBROK | TERRNO, "CPU_ALLOC()");
> +
> +     CPU_ZERO_S(size, mask);
> +     if (sched_getaffinity(0, size, mask) < 0) {
> +             CPU_FREE(mask);
> +             tst_brk(TBROK | TERRNO, "sched_getaffinity()");
> +     }
> +
> +     for (long i = 0; i < ncpus; i++) {
> +             if (CPU_ISSET_S((int)i, size, mask)) {
> +                     cpu = (int)i;
> +                     break;
> +             }
> +     }
> +
> +     if (cpu < 0) {
> +             CPU_FREE(mask);
> +             tst_brk(TBROK, "sched_getaffinity() returned an empty CPU mask");
> +     }
> +
> +     CPU_ZERO_S(size, mask);
> +     CPU_SET_S(cpu, size, mask);
> +     if (sched_setaffinity(0, size, mask) < 0) {
> +             CPU_FREE(mask);
> +             tst_brk(TBROK | TERRNO, "sched_setaffinity()");
> +     }
> +
> +     CPU_FREE(mask);
> +}
> +
> +static void setup(void)
> +{
> +     long online_cpus = count_online_cpus();
> +     long batch = MAX(32L, online_cpus * 2);
> +     long page_size = SAFE_SYSCONF(_SC_PAGESIZE);
> +     long batch_kib = batch * page_size / 1024;
> +
> +     lower_allowance = MAX(20 * 1024L, batch_kib);
> +     if (lower_allowance >= 102400L)
> +             tst_brk(TCONF, "Per-CPU RSS allowance is too large: %li KiB",
> +                     lower_allowance);

As long as we pin to a single CPU the whole batch accounting shouldn't
be needed.

FYI Jan send a similar patch that just pins the process to a single CPU:

https://nam11.safelinks.protection.outlook.com/?url=https%3A%2F%2Fpatchwork.kernel.org%2Fproject%2Fltp%2Fpatch%2Fa308e12198ce96305761a009858fa4a80c91f932.1788518850.git.jstancek%40redhat.com%2F&data=05%7C02%7Cnirmoyd%40nvidia.com%7C5c88fc4e5490430569b008df133e14f3%7C43083d15727340c1b7db39efd9ccc17a%7C0%7C0%7C639250829531773556%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&sdata=qOkWQ7VsVwFS38S1zQgxue5eo8kNqUJ%2BEHl4rN80qIQ%3D&reserved=0<https://patchwork.kernel.org/project/ltp/patch/a308e12198ce96305761a009858fa4a80c91f932.1788518850.git.jstancek@redhat.com/>

As far as I can tell that should be enough to fix the test.

--
Cyril Hrubis
chrubis@suse.cz


More information about the ltp mailing list