Skip to contents

These control how a data-parallel operation cuts its work into chunks once charr_threads() has decided how many threads to use. Neither affects whether threads are used at all, and neither changes any result.

Usage

charr_chunks_per_worker(value = NULL)

charr_min_chunk(value = NULL)

Arguments

value

NULL to query the current setting, or a positive whole number.

Value

The current setting when querying. When setting, the previous value is returned invisibly.

Details

Threads draw chunks one at a time rather than taking a fixed slice each, so a thread that finishes a cheap chunk comes back for another. That is what keeps uneven input balanced: element cost is not uniform, and one long string can outweigh a thousand short ones.

charr_chunks_per_worker() sets how many chunks each thread should get, defaulting to 128. Higher values balance skewed input better and cost a little more bookkeeping. charr_min_chunk() sets the smallest number of elements a chunk may hold, defaulting to 256, which stops a large thread count from cutting the work finer than it is worth handing out.

With n elements and w threads the chunk size is ceiling(n / (w * chunks_per_worker)), raised to at least min_chunk and then lowered if needed so that there are never fewer chunks than threads. A short vector therefore still spreads across every thread.

Like charr_threads(), both settings are held in native code rather than in an R option, because they are only ever read there. These accessors are the only way to change them.

Examples

old <- charr_chunks_per_worker(256)
charr_chunks_per_worker()
#> [1] 256
charr_chunks_per_worker(old)