Architecture
This page follows one annotated function from source to a registered DuckDB function.
Zread documents this repository with generated diagrams — module layering, the registration flow and much more. When this page is not enough, start there.
The pieces
| Crate | Responsibility |
|---|---|
duckfn-macro | Procedural macros. Reads the attributes, validates signatures, and emits the wrapper plus a registration item. |
duckfn | Runtime: adapter traits, value types, inventory registries, entry point glue. |
quack-rs | The DuckDB C API bindings: builders, LogicalType, DataChunk, VectorReader/VectorWriter, SqlMacro, ExtensionError. |
libduckdb-sys | DuckDB's C headers, compiled with the loadable-extension feature. |
duckfn's modules are all pub(crate); the public surface is the set of pub use re-exports in
src/lib.rs, which is why duckfn::DuckOptionResult exists but duckfn::ExtensionError does
not.
1. Expansion
An attribute macro runs common_build, which keeps the original function and adds a module named
after it:
#item_fn // the function, untouched
#vis mod #name { // visibility inherited from the function
use super::*;
#[derive(duckfn::DuckStruct, Debug, Clone, Default)]
#[duck(#attr)] // the attribute arguments are forwarded
pub struct DuckArgsImpl { /* one field per argument */ }
// macro-specific items: the adapter impl and the builders
}
DuckArgsImpl is the bridge between the argument list and the two directions data flows:
#[derive(DuckStruct)] turns it into a DuckColumns type, so the same struct is used to read
arguments from a chunk and to write them back for struct arguments. The adapter impl then calls the
original function. See Attributes for the items each macro emits.
2. Registration
The macro submits a registration item to an inventory registry, which collects items at link time:
pub type DuckRegisterFn = fn(connection: &Connection) -> DuckResult<()>;
pub struct DuckFunctionItem {
pub register_fn: DuckRegisterFn,
}
inventory::collect!(DuckFunctionItem);
There are three registries, one per registration strategy:
| Registry | Collected from | Registered by |
|---|---|---|
DuckFunctionItem | Most macros, and every #[duck_custom_register] function | register_all_duckfn |
DuckAggregateOverloadItem | overloads_name on aggregates | register_all_aggregate_overload |
DuckScalarOverloadItem | overloads_name on scalars | register_all_scalar_overload |
register_all_duckfn applies them in that order, and the overload passes group items by name with
itertools::into_grouping_map_by so that every signature sharing an overloads_name ends up in one
function set. Registry name fields are &'static str rather than String because
inventory::submit! expands into a static initializer, where a String cannot be constructed.
3. Entry point
duckfn_entrypoint!("my_ext");
expands to
quack_rs::entry_point_v2!(my_ext_init_c_api, duckfn::register_all_duckfn);
so the exported symbol is {name}_init_c_api, and the body is a call to register_all_duckfn with
the Connection DuckDB hands over. The name is validated at compile time: it must be non-empty and
contain only lowercase ASCII letters, digits and underscores.
4. Dispatch
With the loadable-extension feature, DuckDB API functions are not linked. Instead they are resolved
through an AtomicPtr table that DuckDB fills in when it loads the extension and calls the entry
point:
This is why the extension does not need a DuckDB build, and why it is tied to the DuckDB version it was compiled against.
5. The adapters
Each registration kind has an adapter trait that turns Rust values into DuckDB's vector-based callbacks.
Scalar
Reads the whole input chunk into a batch of rows (Vec<Option<Args>>, None meaning a NULL row),
hands that batch to apply_batch and writes the results as a batch. apply_batch's default walks
the rows and calls apply_with_null once per row, which keeps the per-row semantics: it returns
Ok(None) when any non-Option argument is NULL, and that is what short-circuits the row.
A batch implementation overrides apply_batch instead (that is what batch = true generates),
receives the whole batch at once and returns Vec<Option<Output>>; Ok(None) from it means the whole
batch is NULL, and the returned length must match the batch size.
null_handling() defaults to DefaultNullHandling and is overridden to SpecialNullHandling by
special_null_handling = true. Likewise volatile() defaults to false and is overridden to true
by volatile = true, which makes registration call duckdb_scalar_function_set_volatile
(DuckDB 1.5+). varargs_element_type() defaults to None; with varargs = true it returns the
element type of the signature's last Vec<T> and registration calls
duckdb_scalar_function_set_varargs (DuckDB 1.5+), while the callback streams the extra columns
through apply_varargs instead of apply (variadic arguments have no stable row structure, so they
are never materialised into a batch).
Aggregate
DuckDB's six callbacks are implemented on the state type:
| Callback | Behaviour |
|---|---|
c_state_size / c_state_init | Allocate and initialise the state (Default). |
c_update | Read a row and call handle_row_with_null; NULL rows are skipped by default. |
c_combine | Merge a partial state into another one, for parallel aggregation. |
c_finalize | Call result() per state and write the vector. A non-zero offset is rejected. |
c_state_destroy | Drop the states. |
Aggregate functions and sets are wrapped in RAII guards (AggregateFunctionGuard,
AggregateFunctionSetGuard) so the DuckDB objects are destroyed even when registration fails midway.
duckfn::DuckfnAggregateFunctionSetBuilder exists because quack-rs's set builder can only set one
return type per set, while each duckfn overload has its own Output.
Table
with_state is the bind step: it reads arguments, decides the result columns through
config_result_columns, and returns the row iterator. scan pulls from that iterator once per chunk.
Both are wrapped in catch_unwind, so a panic in either becomes a query error. The iterator type is
DuckFullIterator<T> = Box<dyn Iterator<Item = DuckOptionResult<T>> + Send>.
Copy
COPY ... TO and COPY ... FROM are both built on the runtime dynamic columns
(src/dynamic) rather than on DuckValueType, because a copy function's columns are
only known during bind.
COPY ... TO is driven through four callbacks: bind turns the output columns into a
DuckResultSchema (one DuckTypeDesc::from_logical_type per column) and reads the copy options,
keeping both as bind data; global_init calls DuckCopyToWriter::open(path, schema, options) and
stores the writer as global state; sink reads each chunk into Vec<DuckDynamicRow> through
DuckDynamicRow::read_batch and calls the annotated function; finalize calls
DuckCopyToWriter::finish.
COPY ... FROM is not a copy-function callback at all: it is an ordinary table function whose scan
produces the rows, wired to the format with duckdb_copy_function_set_copy_from_function. quack-rs
0.16 has no copy_from helper and does not expose the raw table-function handle, so the adapter builds
that table function directly through libduckdb_sys while reusing quack-rs' FfiBindData /
FfiInitData for the bind → init → scan state hand-off. bind parses Args and reads the target
table's schema with duckdb_table_function_bind_get_result_column_* (a COPY FROM reader declares no
result columns); scan calls the annotated batch function and writes the rows through
DuckDynamicRow::write_batch; the reader's finish runs from the init-data destructor, because a table
function has no finalize callback.
Bind data and global state are Boxed and handed to DuckDB with a destructor callback that drops them,
and every callback is wrapped in catch_unwind with errors reported through set_error. This API comes
from DuckDB 1.5.0+, so the modules, the adapters and the quack-rs re-exports are all behind the
duckdb-1-5 feature.
Two ownership traps are worth remembering when touching this code. The strings returned by
duckdb_table_function_bind_get_result_column_name and the value returned by
duckdb_copy_function_bind_get_options are owned by DuckDB and must not be freed — doing so
corrupts the heap (0xC0000374) — so the adapter only borrows them. The reader table-function handle
handed to duckdb_copy_function_set_copy_from_function is deliberately not destroyed: whether
DuckDB copies or takes it over is undocumented, and a double free would be fatal while a per-LOAD
leak is harmless.
Cast
The wrapper receives a count, an input vector and an output vector, and calls the function per row.
Errors are handled according to the cast mode: CastMode::Normal (CAST) fails the query, while
CastMode::Try (TRY_CAST) records a row error and writes NULL.
Replacement scan
scan_callback receives the unresolved table name. handle_info calls the user's handle_path, and
on Some(table_fn) it sets the function to call and adds the path as the first VARCHAR parameter.
Non-UTF-8 names are skipped, and Err is reported through duckdb_replacement_scan_set_error.
SQL macro
The simplest adapter: the returned SqlMacro is registered, or a returned string is executed with
duckdb_query. That is why a string may contain several statements.
6. Value types
DuckValueType is the trait that makes a Rust type usable as an argument or result:
| Direction | Methods |
|---|---|
| Type identity | type_id(), logical_type(), from_null() |
| Reading | create_reader, read_valid, read_slot, read_by_duck_value* |
| Writing | write_batch, write_valid, write_null, write_finish |
Implementors override the *_valid half; read and write are not meant to be overridden.
DuckValueReader and DuckValueWriter carry the vector plus child readers/writers, which is how
nested types recurse into LIST, MAP, ARRAY and STRUCT children. #[derive(DuckStruct)] additionally
generates assert_impl_duck_value_type::<T>() calls for every field, so an unsupported field type is
a compile error rather than a runtime surprise.
Nullability is a property of the type rather than of the container: Option<T> implements
DuckValueType with the same logical type as T and overrides from_null() to Some(None), while
every other type leaves it at the default None. read_slot() — the entry point used for elements,
map entries and struct fields — applies that fallback, which is what lets a single Vec<T> /
[T; N] / IndexMap<K, V> implementation serve both nullable and non-nullable element types, and
why #[derive(DuckStruct)] no longer inspects field types syntactically.
DuckValueReader also carries a liveness token (Arc<ChunkToken>, exposed as a Weak through
alive_weak()), which DuckLazy<T> uses to defer a read: holding a reader is holding proof that its
vector is valid right now, so a deferred value can tell "still inside the callback" from "the chunk is
gone" and report an error instead of dereferencing a stale vector. DuckLazySlot<T> is the consumer
end of that design: it parses the deferred value once into an aggregate state, carries the parsed
result across combine and reads it back in result() — the token itself never leaves the callback.
DuckStructTrait is the generated struct interface, and three blanket impls connect it to the rest of
the system: DuckValueType (usable as a value), DuckColumns (usable as a table function's output
row), and DuckBindArgs (usable as a table function's arguments).
Where to look
| Question | File |
|---|---|
| Which arguments does an attribute accept? | the macro's own file under duckfn-macro/src/ (e.g. scalar_function.rs) |
| What does a macro emit? | that same file plus the shared common.rs; the derives are in duck_struct_derive.rs / duck_enum_derive.rs |
| How is the entry point generated? | duckfn-macro/src/entrypoint.rs |
| How does registration work? | src/register.rs |
| How is a callback implemented? | src/functions/*_adapter.rs |
| How is a type converted? | src/value_types/*.rs |
Next
- Attributes — the user-facing view of expansion.
- Contributing — working on these crates.