Content
previous page: Scaling behavior of psana1 det.calib method in multicore processing with mpi
2024-02-06 Test of milano216 host with perf
Test description
Using command:
perf stat -e cache-references,cache-misses,cycles,instructions,branches,branch-misses,faults,migrations,page-faults,L1-dcache-load-misses,L1-icache-load-misses python test-scaling-subproc.py <parameter>
where parameter defines test for different number of CPUs, e.g. <parameter> = −1,−2,−8,−13,−16,−17,−18 stands for test on single, 8, 16, 32, 56, 64, 128 CPUs.
Results
Summary
number of CPU | cache- references | cache- misses | cycles | instructions | branches | branch- misses | faults | page-faults | L1-dcache- load-misses | L1-icache- load-misses | cmt |
---|---|---|---|---|---|---|---|---|---|---|---|
1 | 4,522,410,200 | 112,207,635 | 2,353,206,592 | 2,169,783,808 | 7,173,374 | ||||||
8 | 35,293,654,947 | 675,772,563 | 18,710,029,709 | 17,164,781,068 | 42,407,266 | ||||||
16 | 71,125,012,043 | 2,509,743,885 | 37,401,077,277 | 34,764,585,133 | 82,908,203 | ||||||
32 | 140,229,421,945 | 5,022,345,750 | 74,783,808,615 | 68,615,480,748 | 163,094,161 | ||||||
56 | 245,664,589,385 | 5,986,128,102 | 130,897,170,304 | 119,933,873,577 | 288,403,921 | ||||||
64 | 281,639,175,978 | 8,968,404,974 | 149,569,155,086 | 137,584,278,754 | 330,750,296 | ||||||
120 | 532,229,037,371 | 14,227,944,434 | 29,404,359,241,173 | 51,095,884,028,391 | 7,053,547,766,317 | 280,479,284,507 | 73,250,012 | 73,250,012 | 260,078,672,869 | 618,858,635 | |
2024-02-07 Test of milano216 host with command perf
Commands
perf stat -e cache-references,cache-misses,cycles,instructions,branches,branch-misses,faults,migrations,page-faults,L1-dcache-load-misses,L1-icache-load-misses,dTLB-load-misses,iTLB-load-misses mpirun -n 1 python Detector/examples/test-scaling-mpi.py
Results
ana-4.0.59-py3 [dubrovin@sdfmilan216:~/LCLS/con-py3]$ perf stat -e cache-references,cache-misses,cycles,instructions,branches,branch-misses,faults,migrations,page-faults,L1-dcache-load-misses,L1-icache-load-misses,dTLB-load-misses,iTLB-load-misses mpirun -n 1 python Detector/examples/test-scaling-mpi.py ... Performance counter stats for 'mpirun -n 1 python Detector/examples/test-scaling-mpi.py': 4,448,830,552 cache-references:u (50.00%) 90,374,312 cache-misses:u # 2.031 % of all cache refs (50.00%) 222,814,516,280 cycles:u (50.02%) 426,700,282,993 instructions:u # 1.92 insn per cycle (50.01%) 58,876,394,584 branches:u (50.01%) 2,343,687,188 branch-misses:u # 3.98% of all branches (50.01%) 635,183 faults:u 0 migrations:u 635,183 page-faults:u 2,158,358,417 L1-dcache-load-misses:u (50.00%) 5,694,036 L1-icache-load-misses:u (49.99%) 4,282,821 dTLB-load-misses:u (49.99%) 890,671 iTLB-load-misses:u (50.00%) 73.297275789 seconds time elapsed 69.795728000 seconds user 2.318007000 seconds sys ana-4.0.59-py3 [dubrovin@sdfmilan216:~/LCLS/con-py3]$ perf stat -e cache-references,cache-misses,cycles,instructions,branches,branch-misses,faults,migrations,page-faults,L1-dcache-load-misses,L1-icache-load-misses,dTLB-load-misses,iTLB-load-misses mpirun -n 80 python Detector/examples/test-scaling-mpi.py ... Performance counter stats for 'mpirun -n 80 python Detector/examples/test-scaling-mpi.py': 349,526,509,383 cache-references:u (50.01%) 5,932,480,814 cache-misses:u # 1.697 % of all cache refs (50.00%) 18,768,444,974,036 cycles:u (50.00%) 33,983,153,714,284 instructions:u # 1.81 insn per cycle (49.99%) 4,684,730,635,234 branches:u (49.99%) 186,649,297,019 branch-misses:u # 3.98% of all branches (50.00%) 52,121,421 faults:u 0 migrations:u 52,121,421 page-faults:u 171,500,392,922 L1-dcache-load-misses:u (50.00%) 267,672,856 L1-icache-load-misses:u (50.00%) 339,145,247 dTLB-load-misses:u (50.01%) 69,780,394 iTLB-load-misses:u (50.01%) 92.952500273 seconds time elapsed 6501.353593000 seconds user 410.844719000 seconds sys