fix(example): batch a buffer sequence in cuda_device_stream::write_some - #392
Merged
sgerbino merged 1 commit intoAug 27, 2026
Merged
Conversation
|
An automated preview of the documentation is available at https://392.capy.prtest3.cppalliance.org/index.html If more commits are pushed to the pull request, the docs will rebuild at the same URL. 2026-08-27 22:03:16 UTC |
sgerbino
force-pushed
the
example/cuda-batched-write
branch
from
August 27, 2026 22:06
3676b0f to
b7dee5a
Compare
write_some copied only the first buffer of the sequence, so a gather write became one memcpy, one host function, and one suspend per buffer, draining the stream between transfers. Enqueue every buffer as its own cudaMemcpyAsync and follow the last with a single cudaLaunchHostFunc: the stream keeps its queue depth and the coroutine suspends once per call, matching what the WriteStream contract already permits. example/cuda/batched-write runs the batched path: three host buffers go through any_write_stream in one write_some, and the device is checked to hold their concatenation. Registered as a ctest. datamovement stays build-only. The datamovement README names the paper by its current title.
sgerbino
force-pushed
the
example/cuda-batched-write
branch
from
August 27, 2026 22:10
b7dee5a to
d45ae3e
Compare
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## develop #392 +/- ##
========================================
Coverage 98.09% 98.09%
========================================
Files 130 130
Lines 6291 6291
========================================
Hits 6171 6171
Misses 120 120
Flags with carried forward coverage won't be shown. Click here to find out more. Continue to review full report in Codecov by Harness.
🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
write_some copied only the first buffer of the sequence, so a gather write became one memcpy, one host function, and one suspend per buffer, draining the stream between transfers. Enqueue every buffer as its own cudaMemcpyAsync and follow the last with a single cudaLaunchHostFunc: the stream keeps its queue depth and the coroutine suspends once per call, matching what the WriteStream contract already permits.
example/cuda/batched-write runs the batched path: three host buffers go through any_write_stream in one write_some, and the device is checked to hold their concatenation. Registered as a ctest. datamovement stays build-only. The datamovement README names the paper by its current title.