Skip to content

OpenBLAS bottlenecks multithreading benefits in symv.c interface at 8 working threads due to memory allocator lock conflict #5589

Description

@BioTurboNick

I have code that heavily makes use of the LAPACK syevr functions via Julia. They're relatively small matrices (at most 15x15). I'm processing chunks of video frames across multiple threads, and each thread will perform millions of these operations. I've set BLAS threads to 1, which I understand to mean that OpenBLAS just uses the parent thread calling it. (Setting it to anything more than 1 tanks performance generally.)

However, what I've found is that no matter what size computer I run on, performance gains stop once I reach 8 working threads; even worsening with many more. Somehow it seems that OpenBLAS, without itself doing multithreaded computation, is interfering with higher-level multithreading?

If I switch to MKL with 1 thread, I see continued performance improvements through 48 CPUs.

I'm willing to poke around at this as much as I can myself, I'm just not sure where to begin. Where might the bottleneck be?

Image

Activity

  1. martin-frbg commented on Jan 2, 2026

    @martin-frbg
    Collaborator

    Is this with 0.3.30 or an older version ? One fairly trivial thing to check would be if there are any excess threads idling. And perhaps running perf on your test code would provide hints to where it gets held up. If you're only ever going to need 1 thread per call, you could try building a single-threaded (USE_THREAD=0 USE_LOCKING=1) version of the library.

  2. BioTurboNick commented on Jan 3, 2026

    @BioTurboNick
    Author

    Thanks for the tips - turns out it might be the settings that Julia uses for the shipped build of OpenBLAS. If I set Julia to build with a locally compiled version, OpenBLAS has performance ~matching MKL with respect to threading.

    The only difference I can see (so far) is that the distributed one is built with TARGET=GENERIC and when the locally built one is built with TARGET= (blank).

    Would you expect that the GENERIC target would behave this way?

    Passed in for the local build that works great:

    BINARY=64 
    USE_THREAD=1 
    GEMM_MULTITHREADING_THRESHOLD=400 
    NUM_THREADS=512 
    NO_AFFINITY=1 
    TARGET= 
    DYNAMIC_ARCH=1 
    INTERFACE64=1 
    

    For the distributed library that shows the issue:

    BINARY=64 
    USE_THREAD=1 
    GEMM_MULTITHREADING_THRESHOLD=400 
    NUM_THREADS=512 
    NO_AFFINITY=1 
    TARGET= 
    DYNAMIC_ARCH=1 
    INTERFACE64=1 
    BUILD_BFLOAT16=1 
    

    This is with v0.3.29 because Julia hasn't shipped a v0.3.30 yet.

  3. BioTurboNick commented on Jan 3, 2026

    @BioTurboNick
    Author

    Nevermind that bit - seems I was on the wrong track. I'll keep investigating.

  4. martin-frbg commented on Jan 4, 2026

    @martin-frbg
    Collaborator

    Yeah, the "TARGET=GENERIC" only ensures that all the common code - thread startup, blas call argument handling, matrix partitioning etc - is built without telling the compiler to use cpu-specific features like AVX2/AVX512 that only some of the targeted cpus would support. I'd be surprised if that caused a major difference in performance compared to a build that allows the compiler full use of the build host's capabilities (but will likely fail to run on older hardware)

  5. BioTurboNick commented on Jan 4, 2026

    @BioTurboNick
    Author

    I finally got a small reproducer in Julia and did some profiling and it seems like the hangup may lie in how BLAS memory allocation is interacting with threads?

    Image
  6. martin-frbg commented on Jan 4, 2026

    @martin-frbg
    Collaborator

    that would be the buffer allocation in interface/symv.c (and similarly done for gemm and a few others) which should not pose a problem unless you're low on memory already. (#4665 argues that this allocation has become redundant due to a later code change and is only wasting memory now, but I haven't gotten around to verifying this and doing any related cleanup)

  7. BioTurboNick commented on Jan 4, 2026

    @BioTurboNick
    Author

    Ah okay. I suppose the issue here is that the different threads are all competing for the memory allocator, which is what's bottlenecking me here?

  8. martin-frbg commented on Jan 4, 2026

    @martin-frbg
    Collaborator

    possibly, yes. but unfortunately I can't simply suggest to throw out the blas_memory_alloc and ...free as I am not immediately sure that identifying the correct buffer id is straightforward at that point in the call graph.

  9. changed the title [-]OpenBLAS bottlenecks multithreading benefits at 8 working threads despite setting BLAS threads to 1[/-] [+]OpenBLAS bottlenecks multithreading benefits in `symv.c` interface at 8 working threads due to memory allocator lock conflict[/+] on Jan 5, 2026
  10. ViralBShah commented on Sep 24, 2026

    @ViralBShah
    Contributor

    I ran into this with benchmarking sparse cholesky performance in Julia when running multiple solves in parallel across multiple threads. The issue is easily reproducible by calling trsv on lots of small matrices (20x20) simultaneously from multiple threads - which is what the CHOLMOD supernodal solve does.

  11. martin-frbg commented on Sep 24, 2026

    @martin-frbg
    Collaborator

    I can probably solve this for trsv as that one is not parallelized (so will always use the main thread and thus could be made to use a permanent buffer allocated there), but I haven't had time to come up with a solution for the general case (exemplified by symv in the issue subject).
    (Which boils down to "how to assign an array of preallocated but not thread-local memory buffers to a bunch of threads we're going to create later, without needing yet another buffer tracking structure that needs locking". Or perhaps preallocating isn't such a good
    idea at all, or the interfaces should each have a static array large enough to serve as buffer for "small" matrices, and malloc only for larger cases ??)

  12. ViralBShah commented on Sep 24, 2026

    @ViralBShah
    Contributor

    Solving trsv will definitely be a huge help in terms of helping SuiteSparse and scaling up CHOLMOD, which otherwise is slowing down 3-4x on 4 threads for me.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions