Repository navigation
openblas_set_num_threads() is silently overridden when built with USE_OPENMP #5806
Description
Activity
When
USE_OPENMPis defined,goto_set_num_threads()should also callomp_set_num_threads(num_threads)so that subsequent calls toomp_get_max_threads()innum_cpu_avail()return the value the user requested.I have not tested this yet, but it seems the most sensical to me that this should work.
I found that this does work indeed:
diff --git a/driver/others/blas_server_omp.c b/driver/others/blas_server_omp.c index 38b48fc84..75cd63dc2 100644 --- a/driver/others/blas_server_omp.c +++ b/driver/others/blas_server_omp.c @@ -111,6 +111,7 @@ void goto_set_num_threads(int num_threads) { } blas_cpu_number = num_threads; + omp_set_num_threads(num_threads); adjust_thread_buffers(); #if defined(ARCH_MIPS64) || defined(ARCH_LOONGARCH64)
But I don't think anymore that it's a good idea to do that, because it sets the global OpenMP thread limit to that, and maybe the user wants to limit only the thread limit of OpenBLAS to get it deterministic, but continue to use OpenMP in their own code that may be parallel and still deterministic.
So it would be better to just make that when the user calls
openblas_set_num_threads(), only OpenBLAS's own OpenMP loops are limited to that number of threads.I am trying to make sense of this whole code but it is really quite difficult.
Consider:
extern int blas_server_avail; extern int blas_omp_number_max; extern int blas_omp_threads_local; static __inline int num_cpu_avail(int level) { #ifdef USE_OPENMP int openmp_nthreads; openmp_nthreads=omp_get_max_threads(); if (omp_in_parallel()) openmp_nthreads = blas_omp_threads_local; #endif #ifndef USE_OPENMP if (blas_cpu_number == 1 #else if (openmp_nthreads == 1 #endif ) return 1; #ifdef USE_OPENMP if (openmp_nthreads > blas_omp_number_max){ #ifdef DEBUG fprintf(stderr,"WARNING - more OpenMP threads requested (%d) than available (%d)\n",openmp_nthreads,blas_omp_number_max); #endif openmp_nthreads = blas_omp_number_max; } if (blas_cpu_number != openmp_nthreads) { goto_set_num_threads(openmp_nthreads); } #endif return blas_cpu_number; }
There are so many questions:
-
What is the purpose of this function? There are no comments, no documentaiton, no intent.
-
The formatting makes it impossible to read.
-
int num_cpu_avail()suggests that it returns numbers of threads to use, a read-only, pure computation. -
Then why does it also have a side effect of calling
goto_set_num_threads()? -
Above it are 3 variables containing state. What is their purpose? Unclear.
-
What is the concrete difference in purpose between
blas_cpu_numberandblas_num_threads, in comparison to those 3 state variables mentioned above such asblas_omp_number_max?- If there are 4 variables all describing thread counts, it would be very helpful if anywhere in the code a comment said what the difference between them is.
-
The USAGE.md it says
In OpenBLAS, we mange a pool of memory buffers and allocate the number of buffers as the following.
#define NUM_BUFFERS (MAX_CPU_NUMBER * 2)But it doesn't explain anywhere why this is made a compile-time constant, given that different computers clearly have different numbers of cores.
Furtherthe setting of NUM_THREADS can be relevant even for a single-threaded build of OpenBLAS
OK, that's confusing, "and what does "can be relevant" mean concretely?
In some cases, the affected code may simply crash or throw a segmentation fault without displaying the above warning first.
That sounds very bad.
But again nowhere does it say: Why?
If this hardcoding ofNUM_THREADSmakes all these problems, why is it done?
Why not just do what every normal program does, and determine the number of threads to use at runtime?
This makes it very difficult to work on even the most trivial fixes, such as making
openblas_set_num_threads(1)make OpenBLAS use 1 thread.-
The
README.mdalso saysWe provide the following functions to control the number of threads at runtime:
void goto_set_num_threads(int num_threads); void openblas_set_num_threads(int num_threads);
Note that these are only used once at library initialization, and are not available for
fine-tuning thread numbers in individual BLAS calls.
If you compile this library withUSE_OPENMP=1, you should use the above functions too.It says
these are only used once at library initialization,
yet in every call to
num_cpu_avail(),goto_set_num_threads()is called, so clearly not only once.It says
If you compile this library with
USE_OPENMP=1, you should use the above functions too.yet, as this issue describes, they don't really have an effect in that case.
PR that should fix it properly: #5808
- added a commit that references this issue
on Jul 12, 2026
Summary
When OpenBLAS is built with
USE_OPENMP, callingopenblas_set_num_threads(1)before a BLAS operation has no lasting effect.The thread count is unconditionally overridden back to
omp_get_max_threads()on every BLAS call, making the API a silent no-op.When a user uses
openblas_set_num_threads(1)in the hope to get deterministic BLAS results, they will eventually find that this does not work, and results are still nondeterministic, because of this.The workaround is to run with the
OMP_NUM_THREADS=1environment variable set. But it seems very wrong that this cannot be done with code, and that a function that is explicitly named to set the number of threads, does not actually do that.Call chain demonstrating the problem
openblas_set_num_threads(1)which callsgoto_set_num_threads(1), settingblas_cpu_number = 1.cblas_sgemm(); code is invoid CNAME()ingemm.c.args.nthreads = get_gemm_optimal_nthreads(MNK)callsnum_cpu_avail(3).num_cpu_avail()callsomp_get_max_threads(), which returns the OpenMP default (all CPUs, sinceOMP_NUM_THREADSis unset).if (blas_cpu_number != openmp_nthreads)is true (1 != N).goto_set_num_threads(openmp_nthreads)overridesblas_cpu_numberback to N.num_cpu_avail()returnsblas_cpu_number(now N), which flows toGEMM_THREAD(..., args.nthreads)then toexec_blas(num_cpu, queue)then to#pragma omp parallel for.Root cause
goto_set_num_threads()does NOT callomp_set_num_threads()-- confirmed by searching the entiredriver/directory (zero results). So the OpenMP runtime's idea of the thread count is never updated, andnum_cpu_avail()always re-syncsblas_cpu_numberfromomp_get_max_threads().Expected behavior
openblas_set_num_threads(1)should durably limit OpenBLAS to 1 thread, even when built withUSE_OPENMP. The only current workaround is settingOMP_NUM_THREADS=1in the environment, which is a process-global side effect affecting all OpenMP consumers.Suggested fix
When
USE_OPENMPis defined,goto_set_num_threads()should also callomp_set_num_threads(num_threads)so that subsequent calls toomp_get_max_threads()innum_cpu_avail()return the value the user requested.I have not tested this yet, but it seems the most sensical to me that this should work.
Practical impact
This breaks deterministic computation for any library that links OpenBLAS and tries to force single-threaded BLAS via the documented API. Multi-threaded floating-point reductions in
exec_blas()produce nondeterministic results due to different summation orders across runs.I found it when trying to make llama-cpp deterministic with its
--threads 1option and found that it didn't work.Version
Tested on commit
8cecf899e(v0.3.32).