Skip to content

c.parallel: device wrappers as code, not format strings - #3439

Merged
griwes merged 12 commits into
NVIDIA:mainfrom
griwes:feature/better-device-wrappers
Apr 21, 2025
Merged

c.parallel: device wrappers as code, not format strings#3439
griwes merged 12 commits into
NVIDIA:mainfrom
griwes:feature/better-device-wrappers

Conversation

@griwes

@griwes griwes commented Jan 17, 2025

Copy link
Copy Markdown
Contributor

Description

So far, even with the recent refactoring, the code for the device wrappers lived in format strings. This has a number of downsides, and this PR aims to resolve those.

The wrappers are now actual C++ templates, can be edited with the help of LSP and code formatters, and attempts to obtain their type names will now do some rudimentary type checking in the host code. This is not perfect, as things like cccl_op_t carry basically no type information, but is an improvement over the status quo.

Also added is a CMake target that bundles all the wrappers into a file that contains their contents, preprocessed to a degree, wrapped into a string, so all the code can be simply added into the NVRTC TU wherever needed.

Resolves #2525
Resolves #2665
Resolves #2918

Checklist

  • New or existing tests cover these changes.
  • The documentation is up to date with these changes.

@griwes griwes self-assigned this Jan 17, 2025
@griwes
griwes requested review from a team as code owners January 17, 2025 17:26
@github-actions

Copy link
Copy Markdown
Contributor
🟩 CI finished in 51m 21s: Pass: 100%/3 | Total: 1h 03m | Avg: 21m 13s | Max: 50m 58s
  • 🟩 cccl_c_parallel: Pass: 100%/2 | Total: 12m 42s | Avg: 6m 21s | Max: 10m 39s

    🟩 cpu
      🟩 amd64              Pass: 100%/2   | Total: 12m 42s | Avg:  6m 21s | Max: 10m 39s
    🟩 ctk
      🟩 12.6               Pass: 100%/2   | Total: 12m 42s | Avg:  6m 21s | Max: 10m 39s
    🟩 cudacxx
      🟩 nvcc12.6           Pass: 100%/2   | Total: 12m 42s | Avg:  6m 21s | Max: 10m 39s
    🟩 cudacxx_family
      🟩 nvcc               Pass: 100%/2   | Total: 12m 42s | Avg:  6m 21s | Max: 10m 39s
    🟩 cxx
      🟩 GCC13              Pass: 100%/2   | Total: 12m 42s | Avg:  6m 21s | Max: 10m 39s
    🟩 cxx_family
      🟩 GCC                Pass: 100%/2   | Total: 12m 42s | Avg:  6m 21s | Max: 10m 39s
    🟩 gpu
      🟩 v100               Pass: 100%/2   | Total: 12m 42s | Avg:  6m 21s | Max: 10m 39s
    🟩 jobs
      🟩 Build              Pass: 100%/1   | Total:  2m 03s | Avg:  2m 03s | Max:  2m 03s
      🟩 Test               Pass: 100%/1   | Total: 10m 39s | Avg: 10m 39s | Max: 10m 39s
    
  • 🟩 python: Pass: 100%/1 | Total: 50m 58s | Avg: 50m 58s | Max: 50m 58s

    🟩 cpu
      🟩 amd64              Pass: 100%/1   | Total: 50m 58s | Avg: 50m 58s | Max: 50m 58s
    🟩 ctk
      🟩 12.6               Pass: 100%/1   | Total: 50m 58s | Avg: 50m 58s | Max: 50m 58s
    🟩 cudacxx
      🟩 nvcc12.6           Pass: 100%/1   | Total: 50m 58s | Avg: 50m 58s | Max: 50m 58s
    🟩 cudacxx_family
      🟩 nvcc               Pass: 100%/1   | Total: 50m 58s | Avg: 50m 58s | Max: 50m 58s
    🟩 cxx
      🟩 GCC13              Pass: 100%/1   | Total: 50m 58s | Avg: 50m 58s | Max: 50m 58s
    🟩 cxx_family
      🟩 GCC                Pass: 100%/1   | Total: 50m 58s | Avg: 50m 58s | Max: 50m 58s
    🟩 gpu
      🟩 v100               Pass: 100%/1   | Total: 50m 58s | Avg: 50m 58s | Max: 50m 58s
    🟩 jobs
      🟩 Test               Pass: 100%/1   | Total: 50m 58s | Avg: 50m 58s | Max: 50m 58s
    

👃 Inspect Changes

Modifications in project?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
CUDA Experimental
python
+/- CCCL C Parallel Library
Catch2Helper

Modifications in project or dependencies?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
CUDA Experimental
+/- python
+/- CCCL C Parallel Library
Catch2Helper

🏃‍ Runner counts (total jobs: 3)

# Runner
2 linux-amd64-gpu-v100-latest-1
1 linux-amd64-cpu16

@shwina shwina left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I can't provide feedback on the implementation, but it's definitely a quality-of-life improvement to be able to define the kernels in code rather than as strings.

@github-actions

github-actions Bot commented Feb 4, 2025

Copy link
Copy Markdown
Contributor
🟩 CI finished in 28m 38s: Pass: 100%/3 | Total: 32m 36s | Avg: 10m 52s | Max: 25m 09s
  • 🟩 cccl_c_parallel: Pass: 100%/2 | Total: 7m 27s | Avg: 3m 43s | Max: 4m 59s

    🟩 cpu
      🟩 amd64              Pass: 100%/2   | Total:  7m 27s | Avg:  3m 43s | Max:  4m 59s
    🟩 ctk
      🟩 12.8               Pass: 100%/2   | Total:  7m 27s | Avg:  3m 43s | Max:  4m 59s
    🟩 cudacxx
      🟩 nvcc12.8           Pass: 100%/2   | Total:  7m 27s | Avg:  3m 43s | Max:  4m 59s
    🟩 cudacxx_family
      🟩 nvcc               Pass: 100%/2   | Total:  7m 27s | Avg:  3m 43s | Max:  4m 59s
    🟩 cxx
      🟩 GCC13              Pass: 100%/2   | Total:  7m 27s | Avg:  3m 43s | Max:  4m 59s
    🟩 cxx_family
      🟩 GCC                Pass: 100%/2   | Total:  7m 27s | Avg:  3m 43s | Max:  4m 59s
    🟩 gpu
      🟩 rtx2080            Pass: 100%/2   | Total:  7m 27s | Avg:  3m 43s | Max:  4m 59s
    🟩 jobs
      🟩 Build              Pass: 100%/1   | Total:  2m 28s | Avg:  2m 28s | Max:  2m 28s
      🟩 Test               Pass: 100%/1   | Total:  4m 59s | Avg:  4m 59s | Max:  4m 59s
    
  • 🟩 python: Pass: 100%/1 | Total: 25m 09s | Avg: 25m 09s | Max: 25m 09s

    🟩 cpu
      🟩 amd64              Pass: 100%/1   | Total: 25m 09s | Avg: 25m 09s | Max: 25m 09s
    🟩 ctk
      🟩 12.8               Pass: 100%/1   | Total: 25m 09s | Avg: 25m 09s | Max: 25m 09s
    🟩 cudacxx
      🟩 nvcc12.8           Pass: 100%/1   | Total: 25m 09s | Avg: 25m 09s | Max: 25m 09s
    🟩 cudacxx_family
      🟩 nvcc               Pass: 100%/1   | Total: 25m 09s | Avg: 25m 09s | Max: 25m 09s
    🟩 cxx
      🟩 GCC13              Pass: 100%/1   | Total: 25m 09s | Avg: 25m 09s | Max: 25m 09s
    🟩 cxx_family
      🟩 GCC                Pass: 100%/1   | Total: 25m 09s | Avg: 25m 09s | Max: 25m 09s
    🟩 gpu
      🟩 rtx2080            Pass: 100%/1   | Total: 25m 09s | Avg: 25m 09s | Max: 25m 09s
    🟩 jobs
      🟩 Test               Pass: 100%/1   | Total: 25m 09s | Avg: 25m 09s | Max: 25m 09s
    

👃 Inspect Changes

Modifications in project?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
CUDA Experimental
python
+/- CCCL C Parallel Library
Catch2Helper

Modifications in project or dependencies?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
CUDA Experimental
+/- python
+/- CCCL C Parallel Library
Catch2Helper

🏃‍ Runner counts (total jobs: 3)

# Runner
2 linux-amd64-gpu-rtx2080-latest-1
1 linux-amd64-cpu16

Comment thread c/parallel/src/jit_templates/README.md Outdated

@gevtushenko gevtushenko left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's quite a chain of indirections

Comment thread c/parallel/src/reduce.cu Outdated
Comment thread c/parallel/src/reduce.cu Outdated
@github-actions

Copy link
Copy Markdown
Contributor
🟩 CI finished in 43m 37s: Pass: 100%/3 | Total: 55m 32s | Avg: 18m 30s | Max: 40m 15s | Hits: 98%/310
  • 🟩 cccl_c_parallel: Pass: 100%/2 | Total: 15m 17s | Avg: 7m 38s | Max: 12m 45s | Hits: 98%/310

    🟩 cpu
      🟩 amd64              Pass: 100%/2   | Total: 15m 17s | Avg:  7m 38s | Max: 12m 45s | Hits:  98%/310   
    🟩 ctk
      🟩 12.8               Pass: 100%/2   | Total: 15m 17s | Avg:  7m 38s | Max: 12m 45s | Hits:  98%/310   
    🟩 cudacxx
      🟩 nvcc12.8           Pass: 100%/2   | Total: 15m 17s | Avg:  7m 38s | Max: 12m 45s | Hits:  98%/310   
    🟩 cudacxx_family
      🟩 nvcc               Pass: 100%/2   | Total: 15m 17s | Avg:  7m 38s | Max: 12m 45s | Hits:  98%/310   
    🟩 cxx
      🟩 GCC13              Pass: 100%/2   | Total: 15m 17s | Avg:  7m 38s | Max: 12m 45s | Hits:  98%/310   
    🟩 cxx_family
      🟩 GCC                Pass: 100%/2   | Total: 15m 17s | Avg:  7m 38s | Max: 12m 45s | Hits:  98%/310   
    🟩 gpu
      🟩 rtx2080            Pass: 100%/2   | Total: 15m 17s | Avg:  7m 38s | Max: 12m 45s | Hits:  98%/310   
    🟩 jobs
      🟩 Build              Pass: 100%/1   | Total:  2m 32s | Avg:  2m 32s | Max:  2m 32s | Hits:  98%/155   
      🟩 Test               Pass: 100%/1   | Total: 12m 45s | Avg: 12m 45s | Max: 12m 45s | Hits:  98%/155   
    
  • 🟩 python: Pass: 100%/1 | Total: 40m 15s | Avg: 40m 15s | Max: 40m 15s

    🟩 cpu
      🟩 amd64              Pass: 100%/1   | Total: 40m 15s | Avg: 40m 15s | Max: 40m 15s
    🟩 ctk
      🟩 12.8               Pass: 100%/1   | Total: 40m 15s | Avg: 40m 15s | Max: 40m 15s
    🟩 cudacxx
      🟩 nvcc12.8           Pass: 100%/1   | Total: 40m 15s | Avg: 40m 15s | Max: 40m 15s
    🟩 cudacxx_family
      🟩 nvcc               Pass: 100%/1   | Total: 40m 15s | Avg: 40m 15s | Max: 40m 15s
    🟩 cxx
      🟩 GCC13              Pass: 100%/1   | Total: 40m 15s | Avg: 40m 15s | Max: 40m 15s
    🟩 cxx_family
      🟩 GCC                Pass: 100%/1   | Total: 40m 15s | Avg: 40m 15s | Max: 40m 15s
    🟩 gpu
      🟩 rtx2080            Pass: 100%/1   | Total: 40m 15s | Avg: 40m 15s | Max: 40m 15s
    🟩 jobs
      🟩 Test               Pass: 100%/1   | Total: 40m 15s | Avg: 40m 15s | Max: 40m 15s
    

👃 Inspect Changes

Modifications in project?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
CUDA Experimental
python
+/- CCCL C Parallel Library
Catch2Helper

Modifications in project or dependencies?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
CUDA Experimental
+/- python
+/- CCCL C Parallel Library
Catch2Helper

🏃‍ Runner counts (total jobs: 3)

# Runner
2 linux-amd64-gpu-rtx2080-latest-1
1 linux-amd64-cpu16

Comment thread c/parallel/src/jit_templates/CMakeLists.txt
Comment thread c/parallel/src/jit_templates/CMakeLists.txt
@github-project-automation github-project-automation Bot moved this from In Review to In Progress in CCCL Mar 5, 2025
@griwes

griwes commented Mar 6, 2025

Copy link
Copy Markdown
Contributor Author

@robertmaynard I believe I addressed your comments.

@github-actions

github-actions Bot commented Mar 6, 2025

Copy link
Copy Markdown
Contributor
🟩 CI finished in 1h 00m: Pass: 100%/3 | Total: 1h 17m | Avg: 25m 58s | Max: 58m 36s | Hits: 98%/310
  • 🟩 cccl_c_parallel: Pass: 100%/2 | Total: 19m 20s | Avg: 9m 40s | Max: 16m 41s | Hits: 98%/310

    🟩 cpu
      🟩 amd64              Pass: 100%/2   | Total: 19m 20s | Avg:  9m 40s | Max: 16m 41s | Hits:  98%/310   
    🟩 ctk
      🟩 12.8               Pass: 100%/2   | Total: 19m 20s | Avg:  9m 40s | Max: 16m 41s | Hits:  98%/310   
    🟩 cudacxx
      🟩 nvcc12.8           Pass: 100%/2   | Total: 19m 20s | Avg:  9m 40s | Max: 16m 41s | Hits:  98%/310   
    🟩 cudacxx_family
      🟩 nvcc               Pass: 100%/2   | Total: 19m 20s | Avg:  9m 40s | Max: 16m 41s | Hits:  98%/310   
    🟩 cxx
      🟩 GCC13              Pass: 100%/2   | Total: 19m 20s | Avg:  9m 40s | Max: 16m 41s | Hits:  98%/310   
    🟩 cxx_family
      🟩 GCC                Pass: 100%/2   | Total: 19m 20s | Avg:  9m 40s | Max: 16m 41s | Hits:  98%/310   
    🟩 gpu
      🟩 rtx2080            Pass: 100%/2   | Total: 19m 20s | Avg:  9m 40s | Max: 16m 41s | Hits:  98%/310   
    🟩 jobs
      🟩 Build              Pass: 100%/1   | Total:  2m 39s | Avg:  2m 39s | Max:  2m 39s | Hits:  98%/155   
      🟩 Test               Pass: 100%/1   | Total: 16m 41s | Avg: 16m 41s | Max: 16m 41s | Hits:  98%/155   
    
  • 🟩 python: Pass: 100%/1 | Total: 58m 36s | Avg: 58m 36s | Max: 58m 36s

    🟩 cpu
      🟩 amd64              Pass: 100%/1   | Total: 58m 36s | Avg: 58m 36s | Max: 58m 36s
    🟩 ctk
      🟩 12.8               Pass: 100%/1   | Total: 58m 36s | Avg: 58m 36s | Max: 58m 36s
    🟩 cudacxx
      🟩 nvcc12.8           Pass: 100%/1   | Total: 58m 36s | Avg: 58m 36s | Max: 58m 36s
    🟩 cudacxx_family
      🟩 nvcc               Pass: 100%/1   | Total: 58m 36s | Avg: 58m 36s | Max: 58m 36s
    🟩 cxx
      🟩 GCC13              Pass: 100%/1   | Total: 58m 36s | Avg: 58m 36s | Max: 58m 36s
    🟩 cxx_family
      🟩 GCC                Pass: 100%/1   | Total: 58m 36s | Avg: 58m 36s | Max: 58m 36s
    🟩 gpu
      🟩 rtx2080            Pass: 100%/1   | Total: 58m 36s | Avg: 58m 36s | Max: 58m 36s
    🟩 jobs
      🟩 Test               Pass: 100%/1   | Total: 58m 36s | Avg: 58m 36s | Max: 58m 36s
    

👃 Inspect Changes

Modifications in project?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
CUDA Experimental
python
+/- CCCL C Parallel Library
Catch2Helper

Modifications in project or dependencies?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
CUDA Experimental
+/- python
+/- CCCL C Parallel Library
Catch2Helper

🏃‍ Runner counts (total jobs: 3)

# Runner
2 linux-amd64-gpu-rtx2080-latest-1
1 linux-amd64-cpu16

Comment thread c/parallel/src/jit_templates/CMakeLists.txt Outdated
@leofang

leofang commented Mar 7, 2025

Copy link
Copy Markdown
Member

cc @NVIDIA/cccl-python-codeowners for vis

@github-actions

github-actions Bot commented Mar 7, 2025

Copy link
Copy Markdown
Contributor
🟩 CI finished in 1h 02m: Pass: 100%/3 | Total: 1h 16m | Avg: 25m 24s | Max: 1h 01m | Hits: 98%/310
  • 🟩 cccl_c_parallel: Pass: 100%/2 | Total: 14m 50s | Avg: 7m 25s | Max: 12m 29s | Hits: 98%/310

    🟩 cpu
      🟩 amd64              Pass: 100%/2   | Total: 14m 50s | Avg:  7m 25s | Max: 12m 29s | Hits:  98%/310   
    🟩 ctk
      🟩 12.8               Pass: 100%/2   | Total: 14m 50s | Avg:  7m 25s | Max: 12m 29s | Hits:  98%/310   
    🟩 cudacxx
      🟩 nvcc12.8           Pass: 100%/2   | Total: 14m 50s | Avg:  7m 25s | Max: 12m 29s | Hits:  98%/310   
    🟩 cudacxx_family
      🟩 nvcc               Pass: 100%/2   | Total: 14m 50s | Avg:  7m 25s | Max: 12m 29s | Hits:  98%/310   
    🟩 cxx
      🟩 GCC13              Pass: 100%/2   | Total: 14m 50s | Avg:  7m 25s | Max: 12m 29s | Hits:  98%/310   
    🟩 cxx_family
      🟩 GCC                Pass: 100%/2   | Total: 14m 50s | Avg:  7m 25s | Max: 12m 29s | Hits:  98%/310   
    🟩 gpu
      🟩 rtx2080            Pass: 100%/2   | Total: 14m 50s | Avg:  7m 25s | Max: 12m 29s | Hits:  98%/310   
    🟩 jobs
      🟩 Build              Pass: 100%/1   | Total:  2m 21s | Avg:  2m 21s | Max:  2m 21s | Hits:  98%/155   
      🟩 Test               Pass: 100%/1   | Total: 12m 29s | Avg: 12m 29s | Max: 12m 29s | Hits:  98%/155   
    
  • 🟩 python: Pass: 100%/1 | Total: 1h 01m | Avg: 1h 01m | Max: 1h 01m

    🟩 cpu
      🟩 amd64              Pass: 100%/1   | Total:  1h 01m | Avg:  1h 01m | Max:  1h 01m
    🟩 ctk
      🟩 12.8               Pass: 100%/1   | Total:  1h 01m | Avg:  1h 01m | Max:  1h 01m
    🟩 cudacxx
      🟩 nvcc12.8           Pass: 100%/1   | Total:  1h 01m | Avg:  1h 01m | Max:  1h 01m
    🟩 cudacxx_family
      🟩 nvcc               Pass: 100%/1   | Total:  1h 01m | Avg:  1h 01m | Max:  1h 01m
    🟩 cxx
      🟩 GCC13              Pass: 100%/1   | Total:  1h 01m | Avg:  1h 01m | Max:  1h 01m
    🟩 cxx_family
      🟩 GCC                Pass: 100%/1   | Total:  1h 01m | Avg:  1h 01m | Max:  1h 01m
    🟩 gpu
      🟩 rtx2080            Pass: 100%/1   | Total:  1h 01m | Avg:  1h 01m | Max:  1h 01m
    🟩 jobs
      🟩 Test               Pass: 100%/1   | Total:  1h 01m | Avg:  1h 01m | Max:  1h 01m
    

👃 Inspect Changes

Modifications in project?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
CUDA Experimental
python
+/- CCCL C Parallel Library
Catch2Helper

Modifications in project or dependencies?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
CUDA Experimental
+/- python
+/- CCCL C Parallel Library
Catch2Helper

🏃‍ Runner counts (total jobs: 3)

# Runner
2 linux-amd64-gpu-rtx2080-latest-1
1 linux-amd64-cpu16

@griwes
griwes requested a review from robertmaynard March 7, 2025 21:39
struct template_id
{};

struct specialization

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could this struct be renamed to specialization_t or specialization_st to make it easier to realize that it describes a type. Ease mental burden of parsing the code.

@rwgk rwgk left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found one trivial grammar error in the Readme.

The code complexity is amazing.

Comment thread c/parallel/src/jit_templates/README.md Outdated
@griwes

griwes commented Mar 18, 2025

Copy link
Copy Markdown
Contributor Author

@robertmaynard can you please re-review?

@github-actions

github-actions Bot commented Apr 7, 2025

Copy link
Copy Markdown
Contributor
🟩 CI finished in 1h 31m: Pass: 100%/3 | Total: 1h 51m | Avg: 37m 00s | Max: 1h 30m | Hits: 96%/330
  • 🟩 cccl_c_parallel: Pass: 100%/2 | Total: 20m 11s | Avg: 10m 05s | Max: 17m 33s | Hits: 96%/330

    🟩 cpu
      🟩 amd64              Pass: 100%/2   | Total: 20m 11s | Avg: 10m 05s | Max: 17m 33s | Hits:  96%/330   
    🟩 ctk
      🟩 12.8               Pass: 100%/2   | Total: 20m 11s | Avg: 10m 05s | Max: 17m 33s | Hits:  96%/330   
    🟩 cudacxx
      🟩 nvcc12.8           Pass: 100%/2   | Total: 20m 11s | Avg: 10m 05s | Max: 17m 33s | Hits:  96%/330   
    🟩 cudacxx_family
      🟩 nvcc               Pass: 100%/2   | Total: 20m 11s | Avg: 10m 05s | Max: 17m 33s | Hits:  96%/330   
    🟩 cxx
      🟩 GCC13              Pass: 100%/2   | Total: 20m 11s | Avg: 10m 05s | Max: 17m 33s | Hits:  96%/330   
    🟩 cxx_family
      🟩 GCC                Pass: 100%/2   | Total: 20m 11s | Avg: 10m 05s | Max: 17m 33s | Hits:  96%/330   
    🟩 gpu
      🟩 rtx2080            Pass: 100%/2   | Total: 20m 11s | Avg: 10m 05s | Max: 17m 33s | Hits:  96%/330   
    🟩 jobs
      🟩 Build              Pass: 100%/1   | Total:  2m 38s | Avg:  2m 38s | Max:  2m 38s | Hits:  93%/165   
      🟩 Test               Pass: 100%/1   | Total: 17m 33s | Avg: 17m 33s | Max: 17m 33s | Hits:  98%/165   
    
  • 🟩 python: Pass: 100%/1 | Total: 1h 30m | Avg: 1h 30m | Max: 1h 30m

    🟩 cpu
      🟩 amd64              Pass: 100%/1   | Total:  1h 30m | Avg:  1h 30m | Max:  1h 30m
    🟩 ctk
      🟩 12.8               Pass: 100%/1   | Total:  1h 30m | Avg:  1h 30m | Max:  1h 30m
    🟩 cudacxx
      🟩 nvcc12.8           Pass: 100%/1   | Total:  1h 30m | Avg:  1h 30m | Max:  1h 30m
    🟩 cudacxx_family
      🟩 nvcc               Pass: 100%/1   | Total:  1h 30m | Avg:  1h 30m | Max:  1h 30m
    🟩 cxx
      🟩 GCC13              Pass: 100%/1   | Total:  1h 30m | Avg:  1h 30m | Max:  1h 30m
    🟩 cxx_family
      🟩 GCC                Pass: 100%/1   | Total:  1h 30m | Avg:  1h 30m | Max:  1h 30m
    🟩 gpu
      🟩 rtx2080            Pass: 100%/1   | Total:  1h 30m | Avg:  1h 30m | Max:  1h 30m
    🟩 jobs
      🟩 Test               Pass: 100%/1   | Total:  1h 30m | Avg:  1h 30m | Max:  1h 30m
    

👃 Inspect Changes

Modifications in project?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
CUDA Experimental
stdpar
python
+/- CCCL C Parallel Library
Catch2Helper

Modifications in project or dependencies?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
CUDA Experimental
stdpar
+/- python
+/- CCCL C Parallel Library
Catch2Helper

🏃‍ Runner counts (total jobs: 3)

# Runner
2 linux-amd64-gpu-rtx2080-latest-1
1 linux-amd64-cpu16

@griwes

griwes commented Apr 9, 2025

Copy link
Copy Markdown
Contributor Author

@NVIDIA/cccl-cmake-codeowners ping for a re-review

@jrhemstad jrhemstad moved this from In Progress to In Review in CCCL Apr 9, 2025
@jrhemstad
jrhemstad requested a review from alliepiper April 9, 2025 16:36

@alliepiper alliepiper left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Only reviewed CMake changes.

The features used here aren't an area of CMake I'm very familiar with, but based on @robertmaynard's comments and my reading of the relevant docs this LGTM.

Approving.

@griwes
griwes enabled auto-merge (squash) April 21, 2025 18:37
@github-actions

Copy link
Copy Markdown
Contributor
🟩 CI finished in 41m 18s: Pass: 100%/5 | Total: 1h 07m | Avg: 13m 32s | Max: 36m 37s | Hits: 98%/326
  • 🟩 python: Pass: 100%/3 | Total: 28m 26s | Avg: 9m 28s | Max: 19m 04s

    🟩 cpu
      🟩 amd64              Pass: 100%/3   | Total: 28m 26s | Avg:  9m 28s | Max: 19m 04s
    🟩 ctk
      🟩 12.8               Pass: 100%/3   | Total: 28m 26s | Avg:  9m 28s | Max: 19m 04s
    🟩 cudacxx
      🟩 nvcc12.8           Pass: 100%/3   | Total: 28m 26s | Avg:  9m 28s | Max: 19m 04s
    🟩 cudacxx_family
      🟩 nvcc               Pass: 100%/3   | Total: 28m 26s | Avg:  9m 28s | Max: 19m 04s
    🟩 cxx
      🟩 GCC13              Pass: 100%/3   | Total: 28m 26s | Avg:  9m 28s | Max: 19m 04s
    🟩 cxx_family
      🟩 GCC                Pass: 100%/3   | Total: 28m 26s | Avg:  9m 28s | Max: 19m 04s
    🟩 gpu
      🟩 rtx2080            Pass: 100%/3   | Total: 28m 26s | Avg:  9m 28s | Max: 19m 04s
    🟩 jobs
      🟩 cuda.cccl          Pass: 100%/1   | Total:  2m 49s | Avg:  2m 49s | Max:  2m 49s
      🟩 cuda.cooperative   Pass: 100%/1   | Total: 19m 04s | Avg: 19m 04s | Max: 19m 04s
      🟩 cuda.parallel      Pass: 100%/1   | Total:  6m 33s | Avg:  6m 33s | Max:  6m 33s
    
  • 🟩 cccl_c_parallel: Pass: 100%/2 | Total: 39m 17s | Avg: 19m 38s | Max: 36m 37s | Hits: 98%/326

    🟩 cpu
      🟩 amd64              Pass: 100%/2   | Total: 39m 17s | Avg: 19m 38s | Max: 36m 37s | Hits:  98%/326   
    🟩 ctk
      🟩 12.8               Pass: 100%/2   | Total: 39m 17s | Avg: 19m 38s | Max: 36m 37s | Hits:  98%/326   
    🟩 cudacxx
      🟩 nvcc12.8           Pass: 100%/2   | Total: 39m 17s | Avg: 19m 38s | Max: 36m 37s | Hits:  98%/326   
    🟩 cudacxx_family
      🟩 nvcc               Pass: 100%/2   | Total: 39m 17s | Avg: 19m 38s | Max: 36m 37s | Hits:  98%/326   
    🟩 cxx
      🟩 GCC13              Pass: 100%/2   | Total: 39m 17s | Avg: 19m 38s | Max: 36m 37s | Hits:  98%/326   
    🟩 cxx_family
      🟩 GCC                Pass: 100%/2   | Total: 39m 17s | Avg: 19m 38s | Max: 36m 37s | Hits:  98%/326   
    🟩 gpu
      🟩 rtx2080            Pass: 100%/2   | Total: 39m 17s | Avg: 19m 38s | Max: 36m 37s | Hits:  98%/326   
    🟩 jobs
      🟩 Build              Pass: 100%/1   | Total:  2m 40s | Avg:  2m 40s | Max:  2m 40s | Hits:  98%/163   
      🟩 Test               Pass: 100%/1   | Total: 36m 37s | Avg: 36m 37s | Max: 36m 37s | Hits:  98%/163   
    

👃 Inspect Changes

Modifications in project?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
CUDA Experimental
stdpar
python
+/- CCCL C Parallel Library
Catch2Helper

Modifications in project or dependencies?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
CUDA Experimental
stdpar
+/- python
+/- CCCL C Parallel Library
Catch2Helper

🏃‍ Runner counts (total jobs: 5)

# Runner
4 linux-amd64-gpu-rtx2080-latest-1
1 linux-amd64-cpu16

@github-actions

Copy link
Copy Markdown
Contributor
🟩 CI finished in 40m 19s: Pass: 100%/5 | Total: 1h 04m | Avg: 12m 56s | Max: 32m 28s | Hits: 55%/326
  • 🟩 python: Pass: 100%/3 | Total: 28m 48s | Avg: 9m 36s | Max: 16m 12s

    🟩 cpu
      🟩 amd64              Pass: 100%/3   | Total: 28m 48s | Avg:  9m 36s | Max: 16m 12s
    🟩 ctk
      🟩 12.8               Pass: 100%/3   | Total: 28m 48s | Avg:  9m 36s | Max: 16m 12s
    🟩 cudacxx
      🟩 nvcc12.8           Pass: 100%/3   | Total: 28m 48s | Avg:  9m 36s | Max: 16m 12s
    🟩 cudacxx_family
      🟩 nvcc               Pass: 100%/3   | Total: 28m 48s | Avg:  9m 36s | Max: 16m 12s
    🟩 cxx
      🟩 GCC13              Pass: 100%/3   | Total: 28m 48s | Avg:  9m 36s | Max: 16m 12s
    🟩 cxx_family
      🟩 GCC                Pass: 100%/3   | Total: 28m 48s | Avg:  9m 36s | Max: 16m 12s
    🟩 gpu
      🟩 rtx2080            Pass: 100%/3   | Total: 28m 48s | Avg:  9m 36s | Max: 16m 12s
    🟩 jobs
      🟩 cuda.cccl          Pass: 100%/1   | Total:  5m 53s | Avg:  5m 53s | Max:  5m 53s
      🟩 cuda.cooperative   Pass: 100%/1   | Total: 16m 12s | Avg: 16m 12s | Max: 16m 12s
      🟩 cuda.parallel      Pass: 100%/1   | Total:  6m 43s | Avg:  6m 43s | Max:  6m 43s
    
  • 🟩 cccl_c_parallel: Pass: 100%/2 | Total: 35m 55s | Avg: 17m 57s | Max: 32m 28s | Hits: 55%/326

    🟩 cpu
      🟩 amd64              Pass: 100%/2   | Total: 35m 55s | Avg: 17m 57s | Max: 32m 28s | Hits:  55%/326   
    🟩 ctk
      🟩 12.8               Pass: 100%/2   | Total: 35m 55s | Avg: 17m 57s | Max: 32m 28s | Hits:  55%/326   
    🟩 cudacxx
      🟩 nvcc12.8           Pass: 100%/2   | Total: 35m 55s | Avg: 17m 57s | Max: 32m 28s | Hits:  55%/326   
    🟩 cudacxx_family
      🟩 nvcc               Pass: 100%/2   | Total: 35m 55s | Avg: 17m 57s | Max: 32m 28s | Hits:  55%/326   
    🟩 cxx
      🟩 GCC13              Pass: 100%/2   | Total: 35m 55s | Avg: 17m 57s | Max: 32m 28s | Hits:  55%/326   
    🟩 cxx_family
      🟩 GCC                Pass: 100%/2   | Total: 35m 55s | Avg: 17m 57s | Max: 32m 28s | Hits:  55%/326   
    🟩 gpu
      🟩 rtx2080            Pass: 100%/2   | Total: 35m 55s | Avg: 17m 57s | Max: 32m 28s | Hits:  55%/326   
    🟩 jobs
      🟩 Build              Pass: 100%/1   | Total:  3m 27s | Avg:  3m 27s | Max:  3m 27s | Hits:  11%/163   
      🟩 Test               Pass: 100%/1   | Total: 32m 28s | Avg: 32m 28s | Max: 32m 28s | Hits:  98%/163   
    

👃 Inspect Changes

Modifications in project?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
CUDA Experimental
stdpar
python
+/- CCCL C Parallel Library
Catch2Helper

Modifications in project or dependencies?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
CUDA Experimental
stdpar
+/- python
+/- CCCL C Parallel Library
Catch2Helper

🏃‍ Runner counts (total jobs: 5)

# Runner
4 linux-amd64-gpu-rtx2080-latest-1
1 linux-amd64-cpu16

@griwes
griwes merged commit 9f254d5 into NVIDIA:main Apr 21, 2025
@github-project-automation github-project-automation Bot moved this from In Review to Done in CCCL Apr 21, 2025
@griwes
griwes deleted the feature/better-device-wrappers branch April 21, 2025 20:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Archived in project

10 participants