Skip to main content

Architecture

This page follows one annotated function from source to a registered DuckDB function.

Diagrams and a deeper walkthrough

Zread documents this repository with generated diagrams — module layering, the registration flow and much more. When this page is not enough, start there.

The pieces​

CrateResponsibility
duckfn-macroProcedural macros. Reads the attributes, validates signatures, and emits the wrapper plus a registration item.
duckfnRuntime: adapter traits, value types, inventory registries, entry point glue.
quack-rsThe DuckDB C API bindings: builders, LogicalType, DataChunk, VectorReader/VectorWriter, SqlMacro, ExtensionError.
libduckdb-sysDuckDB's C headers, compiled with the loadable-extension feature.

duckfn's modules are all pub(crate); the public surface is the set of pub use re-exports in src/lib.rs, which is why duckfn::DuckOptionResult exists but duckfn::ExtensionError does not.

1. Expansion​

An attribute macro runs common_build, which keeps the original function and adds a module named after it:

#item_fn // the function, untouched
#vis mod #name { // visibility inherited from the function
use super::*;
#[derive(duckfn::DuckStruct, Debug, Clone, Default)]
#[duck(#attr)] // the attribute arguments are forwarded
pub struct DuckArgsImpl { /* one field per argument */ }
// macro-specific items: the adapter impl and the builders
}

DuckArgsImpl is the bridge between the argument list and the two directions data flows: #[derive(DuckStruct)] turns it into a DuckColumns type, so the same struct is used to read arguments from a chunk and to write them back for struct arguments. The adapter impl then calls the original function. See Attributes for the items each macro emits.

2. Registration​

The macro submits a registration item to an inventory registry, which collects items at link time:

pub type DuckRegisterFn = fn(connection: &Connection) -> DuckResult<()>;

pub struct DuckFunctionItem {
pub register_fn: DuckRegisterFn,
}

inventory::collect!(DuckFunctionItem);

There are three registries, one per registration strategy:

RegistryCollected fromRegistered by
DuckFunctionItemMost macros, and every #[duck_custom_register] functionregister_all_duckfn
DuckAggregateOverloadItemoverloads_name on aggregatesregister_all_aggregate_overload
DuckScalarOverloadItemoverloads_name on scalarsregister_all_scalar_overload

register_all_duckfn applies them in that order, and the overload passes group items by name with itertools::into_grouping_map_by so that every signature sharing an overloads_name ends up in one function set. Registry name fields are &'static str rather than String because inventory::submit! expands into a static initializer, where a String cannot be constructed.

3. Entry point​

duckfn_entrypoint!("my_ext");

expands to

quack_rs::entry_point_v2!(my_ext_init_c_api, duckfn::register_all_duckfn);

so the exported symbol is {name}_init_c_api, and the body is a call to register_all_duckfn with the Connection DuckDB hands over. The name is validated at compile time: it must be non-empty and contain only lowercase ASCII letters, digits and underscores.

4. Dispatch​

With the loadable-extension feature, DuckDB API functions are not linked. Instead they are resolved through an AtomicPtr table that DuckDB fills in when it loads the extension and calls the entry point:

This is why the extension does not need a DuckDB build, and why it is tied to the DuckDB version it was compiled against.

5. The adapters​

Each registration kind has an adapter trait that turns Rust values into DuckDB's vector-based callbacks.

Scalar​

Reads the whole input chunk into a batch of rows (Vec<Option<Args>>, None meaning a NULL row), hands that batch to apply_batch and writes the results as a batch. apply_batch's default walks the rows and calls apply_with_null once per row, which keeps the per-row semantics: it returns Ok(None) when any non-Option argument is NULL, and that is what short-circuits the row. A batch implementation overrides apply_batch instead (that is what batch = true generates), receives the whole batch at once and returns Vec<Option<Output>>; Ok(None) from it means the whole batch is NULL, and the returned length must match the batch size.

null_handling() defaults to DefaultNullHandling and is overridden to SpecialNullHandling by special_null_handling = true. Likewise volatile() defaults to false and is overridden to true by volatile = true, which makes registration call duckdb_scalar_function_set_volatile (DuckDB 1.5+). varargs_element_type() defaults to None; with varargs = true it returns the element type of the signature's last Vec<T> and registration calls duckdb_scalar_function_set_varargs (DuckDB 1.5+), while the callback streams the extra columns through apply_varargs instead of apply (variadic arguments have no stable row structure, so they are never materialised into a batch).

Aggregate​

DuckDB's six callbacks are implemented on the state type:

CallbackBehaviour
c_state_size / c_state_initAllocate and initialise the state (Default).
c_updateRead a row and call handle_row_with_null; NULL rows are skipped by default.
c_combineMerge a partial state into another one, for parallel aggregation.
c_finalizeCall result() per state and write the vector. A non-zero offset is rejected.
c_state_destroyDrop the states.

Aggregate functions and sets are wrapped in RAII guards (AggregateFunctionGuard, AggregateFunctionSetGuard) so the DuckDB objects are destroyed even when registration fails midway. duckfn::DuckfnAggregateFunctionSetBuilder exists because quack-rs's set builder can only set one return type per set, while each duckfn overload has its own Output.

Table​

with_state is the bind step: it reads arguments, decides the result columns through config_result_columns, and returns the row iterator. scan pulls from that iterator once per chunk. Both are wrapped in catch_unwind, so a panic in either becomes a query error. The iterator type is DuckFullIterator<T> = Box<dyn Iterator<Item = DuckOptionResult<T>> + Send>.

Copy​

COPY ... TO and COPY ... FROM are both built on the runtime dynamic columns (src/dynamic) rather than on DuckValueType, because a copy function's columns are only known during bind.

COPY ... TO is driven through four callbacks: bind turns the output columns into a DuckResultSchema (one DuckTypeDesc::from_logical_type per column) and reads the copy options, keeping both as bind data; global_init calls DuckCopyToWriter::open(path, schema, options) and stores the writer as global state; sink reads each chunk into Vec<DuckDynamicRow> through DuckDynamicRow::read_batch and calls the annotated function; finalize calls DuckCopyToWriter::finish.

COPY ... FROM is not a copy-function callback at all: it is an ordinary table function whose scan produces the rows, wired to the format with duckdb_copy_function_set_copy_from_function. quack-rs 0.16 has no copy_from helper and does not expose the raw table-function handle, so the adapter builds that table function directly through libduckdb_sys while reusing quack-rs' FfiBindData / FfiInitData for the bind → init → scan state hand-off. bind parses Args and reads the target table's schema with duckdb_table_function_bind_get_result_column_* (a COPY FROM reader declares no result columns); scan calls the annotated batch function and writes the rows through DuckDynamicRow::write_batch; the reader's finish runs from the init-data destructor, because a table function has no finalize callback.

Bind data and global state are Boxed and handed to DuckDB with a destructor callback that drops them, and every callback is wrapped in catch_unwind with errors reported through set_error. This API comes from DuckDB 1.5.0+, so the modules, the adapters and the quack-rs re-exports are all behind the duckdb-1-5 feature.

Two ownership traps are worth remembering when touching this code. The strings returned by duckdb_table_function_bind_get_result_column_name and the value returned by duckdb_copy_function_bind_get_options are owned by DuckDB and must not be freed — doing so corrupts the heap (0xC0000374) — so the adapter only borrows them. The reader table-function handle handed to duckdb_copy_function_set_copy_from_function is deliberately not destroyed: whether DuckDB copies or takes it over is undocumented, and a double free would be fatal while a per-LOAD leak is harmless.

Cast​

The wrapper receives a count, an input vector and an output vector, and calls the function per row. Errors are handled according to the cast mode: CastMode::Normal (CAST) fails the query, while CastMode::Try (TRY_CAST) records a row error and writes NULL.

Replacement scan​

scan_callback receives the unresolved table name. handle_info calls the user's handle_path, and on Some(table_fn) it sets the function to call and adds the path as the first VARCHAR parameter. Non-UTF-8 names are skipped, and Err is reported through duckdb_replacement_scan_set_error.

SQL macro​

The simplest adapter: the returned SqlMacro is registered, or a returned string is executed with duckdb_query. That is why a string may contain several statements.

6. Value types​

DuckValueType is the trait that makes a Rust type usable as an argument or result:

DirectionMethods
Type identitytype_id(), logical_type(), from_null()
Readingcreate_reader, read_valid, read_slot, read_by_duck_value*
Writingwrite_batch, write_valid, write_null, write_finish

Implementors override the *_valid half; read and write are not meant to be overridden. DuckValueReader and DuckValueWriter carry the vector plus child readers/writers, which is how nested types recurse into LIST, MAP, ARRAY and STRUCT children. #[derive(DuckStruct)] additionally generates assert_impl_duck_value_type::<T>() calls for every field, so an unsupported field type is a compile error rather than a runtime surprise.

Nullability is a property of the type rather than of the container: Option<T> implements DuckValueType with the same logical type as T and overrides from_null() to Some(None), while every other type leaves it at the default None. read_slot() — the entry point used for elements, map entries and struct fields — applies that fallback, which is what lets a single Vec<T> / [T; N] / IndexMap<K, V> implementation serve both nullable and non-nullable element types, and why #[derive(DuckStruct)] no longer inspects field types syntactically.

DuckValueReader also carries a liveness token (Arc<ChunkToken>, exposed as a Weak through alive_weak()), which DuckLazy<T> uses to defer a read: holding a reader is holding proof that its vector is valid right now, so a deferred value can tell "still inside the callback" from "the chunk is gone" and report an error instead of dereferencing a stale vector. DuckLazySlot<T> is the consumer end of that design: it parses the deferred value once into an aggregate state, carries the parsed result across combine and reads it back in result() — the token itself never leaves the callback.

DuckStructTrait is the generated struct interface, and three blanket impls connect it to the rest of the system: DuckValueType (usable as a value), DuckColumns (usable as a table function's output row), and DuckBindArgs (usable as a table function's arguments).

Where to look​

QuestionFile
Which arguments does an attribute accept?the macro's own file under duckfn-macro/src/ (e.g. scalar_function.rs)
What does a macro emit?that same file plus the shared common.rs; the derives are in duck_struct_derive.rs / duck_enum_derive.rs
How is the entry point generated?duckfn-macro/src/entrypoint.rs
How does registration work?src/register.rs
How is a callback implemented?src/functions/*_adapter.rs
How is a type converted?src/value_types/*.rs

Next​