For this month’s article, I wanted to focus on something near and dear to me - macros.
I’m only exaggerating a bit when I say that. The truth is I just like code that writes other code - generic programming, code generators, macros, and similar techniques.
In C, macros are pretty easy to learn because you can abstract them as simple “text substitution”. But the simplicity of their implementation can often lead to surprising and counter-intuitive results. As such, they are difficult to master.
In Rust, however, macros work slightly differently. They are recognized as a part of the compiler’s parsing process. That can make them do some powerful things, but it also can make them a little hard to learn and understand initially.
Since macros are a huge topic, there’s no way I’m going to be able to cover all of it in one article. So we’re only going to focus on a brief introduction to macros, as well as defining and using declarative macros.
In Rust, there are two different macro systems.
- Declarative Macros (aka “Macros By Example”, aka “
macro_rules!Macros”) - Procedural Macros (often shortened to just “proc-macros”)
Additionally, there are three invocation forms for macros. Of these invocation forms, you must select one to implement.
| Invocation Form | Example | Declarative | Procedural |
|---|---|---|---|
| Function-Like | some_macro!(a, b, c, ...) | ✅ | ✅ |
| Derive | #[derive(MyDeriveMacro)] | ❌ | ✅ |
| Attribute | #[my_macro(a, b, c, ...)] | ❌ | ✅ |
If you want to make something like a derive macro, or an attribute macro - you have no choice, and must use procedural macros to implement them.
However, function-like macros can be implemented either by declarative macros or procedural macros.
I’ve heard comments in the past from folks who consider proc-macros to be more difficult to maintain than declarative macros. I strongly disagree with that sentiment. Though I do believe that not carefully developed proc-macros can be difficult to maintain - but that’s true of any code, really.
However, that doesn’t mean that a proc-macro is always the right tool for the job. You should know about both implementations and when to select one versus the other.
Pros:
- Requires very little setup to get started.
- By default, necessitates no additional dependencies.
- Convenient for some predefined parsing and expansion rules (
ident,literal, etc).
Cons:
- Complex matching and expansion logic can become difficult to read and maintain.
- Less parsing and expansion flexibility than procedural macros.
- Not available at all for derive or attribute invocation forms.
Brief Description:
If you’re coming from a C/C++ background, these are the closest things to what you know as macros.
However, they work quite differently from normal C-style macros because of which part of the compilation process they inject themselves into. Still, they solve a similar kind of problem, and in a close enough way that you could start by mentally slotting them into that camp.
Later sections of the article will continue to expand on the differences.
Pros:
- Can implement all three of the invocation forms (function-like, derive, and attribute).
- Supports custom parsing, allowing scenarios that cannot be expressed using predefined fragment specifiers.
- Can be easier to organize and maintain for more complex macro implementations.
Cons:
- Must be defined in a separate crate with the
proc-macrocrate type. - Requires more setup and a deeper understanding of Rust token streams.
- Often requires helper crates like
synandquoteto make them easier to write, which may increase compile times.
Brief Description:
The way I like to describe these (and why they take a bit more setup than declarative macros) is to say that they’re kind of like compiler extensions. You have more control, but at the cost of needing to interact with macros conceptually at a lower level.
You author a proc-macro crate that exports the macro implementation, and then the compiler can load and invoke the macro when explicitly named. This is why there must be a separate crate, because in order to use the extension it must have already been compiled before the compiler can compile other code that uses the macro.
This is a rough mental model, but it’s a simple way to describe the need for the separate crate and build step.
Before we can explain the difference between C-style and Rust macros, it’s worth it to have a brief refresher on the relevant phases of compiling a program.
- Lexing - Taking raw text and turning it into tokens for further parsing logic to process.
So for instance, this is taking some text like “123 + abc” and converting it into something a bit more structured. A toy example might be something like [Literal(123), Punctuation('+'), Identifier("abc")]. It’s best to have an initial pass at the data to classify the inputs, and then the parsing logic can work with specific tokens instead of just raw text.
- Parsing - Converting the stream of tokens to an abstract syntax tree
In this phase, we organize the tokens into something valid given the language’s grammar rules. With the same example above, we may be in a position to understand that we’re parsing something that must be an expression, and that a valid expression can be formed by combining the tokens [Literal(123), Punctuation('+'), Identifier("abc")] into a structure of Expression::Add(Left=Literal(123), Right=Identifier("abc")).
- Semantic Analysis - Checking the structure to make sure all of the types are correct and everything works in the positions they’re in.
We may be able to form a syntax tree for 123 + abc, but for example if abc is a variable binding containing a string, or there’s no identifier named abc, semantically it won’t make sense. This is what we’re dealing with in the semantic analysis phase. We’re checking name resolution, types, traits, and other things to make sure the usage of the syntax is valid.
There’s more phases to compilation after this, but these are the relevant ones for talking about Rust macros - so let’s stop here for now.
C-style macros are pretty conceptually simple, as such you don’t usually need to understand the phases of compilation to understand how they work. You define a macro, and then the text at the locations where you invoke the macro is replaced with whatever the macro says it should be replaced it with.
This could be imagined as occurring before the compiler’s Lexing Phase.
In C, we expand macros and then continue the tokenization process post-expansion. This is why it’s possible to do some scary things with C-style macros, like open new scopes without closing them, or reference variables defined outside of our macro definition. These are conceptually simple expansion rules, but easy to misuse if you aren’t careful, and can produce surprising and unexpected results.
Rust macros, in comparison, can be abstracted as happening during the Parsing Phase.
In Rust, macros don’t operate on raw text like C macros do. But they don’t operate on a fully-formed abstract syntax tree either. Instead, they operate on the in-between, a sequence of trees of tokens. As the compiler parses the code, it recognizes macro invocations, matches them against specific token patterns (fragments), then expands into a new set of tokens which merges into the surrounding AST.
The following is a deliberately simplified model to demonstrate this. The way the compiler works is actually much more involved and complex than this.
| Phase | Content |
|---|---|
| 1. Raw Code | 123 + foo!(a $ 5) |
| 2. Tokens | [Lit(123), '+', Ident("foo"), '!', Tree([Ident("a"), '$', Lit(5)])] |
| 3. Un-Expanded AST | Expr::Add(Left=Lit(123), Right=Macro(Name="foo", args=[Ident("a"), '$', Lit(5)])) |
| 4. Expanded AST | Expr::Add(Left=Lit(123), Right=MACRO_OUTPUT) |
Knowing this allows us to make the following observations:
- The input provided to a macro must consist of lexable Rust tokens.
- The tokens can be ordered in ways for which there are no grammar rules in Rust (e.g.
a $ 5is a valid sequence of tokens, but not any current valid Rust grammar). - Token groups (like
(a, b, ...),[a, b, ...], or{a, b, ...}) must be balanced - you can’t have an open bracket without a closed bracket. - Macro invocations are only allowed in syntactically supported positions, so there are limits where you can use them.
- For example, we cannot have a macro that transcribes to
+ 5, and then usea foo!()to producea + 5.
- For example, we cannot have a macro that transcribes to
- The expanded fragment must be valid at the position it’s invoked at.
- Similarly,
a + foo!()must return some expression that can slot into being the right-hand argument of an addition operation.
- Similarly,
The last point has an important consequence: A macro invocation in expression position must produce an expression, but in a type position it must produce a type, and so on.
Expression-position invocations thus behave like a single expression for purposes of parsing and operator precedence. For example:
macro_rules! example {
() => { 1 + 3 };
}
fn main() {
// This is equivalent to 5 * (1 + 3), NOT 5 * 1 + 3.
assert_eq!(5 * example!(), 20);
}
Macros are said to be hygienic when identifiers within the macro do not accidentally conflict with, or capture, identifiers in the code that invokes it.
C-style macros are generally unhygienic. They allow you to perform text or token-based substitution without separating identifier contexts. As such, you can refer to identifiers outside of the macro, or even from within the macro from the outside.
If all macros in Rust were fully-hygienic, that would be kind of annoying. Imagine being unable to access a type or function without first providing it to the macro - the system would be rigid and difficult to work with.
As such, Rust employs what’s known as mixed-site hygiene, where some things are resolved at the invocation-site (functions, types, etc), and other things are resolved at the definition-site (loop labels, block labels, local variables; source).
For declarative macros, the place where you’ll usually run into definition-site items causing trouble are local variables.
// A simple macro that expands to a print statement of some `variable`.
//
// I know we haven't gotten to how to write declarative macros yet,
// but simply put this is a macro that takes no "tokens", and expands
// to the tokens inside the curly brackets (`{}`).
macro_rules! print_variable {
() => {
println!("{}", variable)
};
}
fn main() {
// If the macro were unhygienic, this would compile.
//
// However, the macro attempts to resolve `variable` from within the
// context of the macro only - so it cannot access the local defined
// just before the macro invocation.
//
// You will get an error due to this at compile-time.
let variable = "Hello, Macros!";
print_variable!();
}
The fix for this is to provide a way to look-up the variable within our macro definition itself. You can accomplish this by passing the name of the identifier into the macro via some metavariable.
This is NOT the same as passing the value into the macro, you’re just passing the name. When substituted into the expansion, it retains the caller’s context, so variable is resolved as the local variable in main. How the resolved identifier is used in the end simply depends on what the macro expands to.
// Okay now we're saying this macro takes exactly 1 fragment - an identifier.
//
// We then use the identifier in the expansion in the `println` call.
macro_rules! print_variable {
($name:ident) => {
println!("{}", $name)
};
}
fn main() {
// Then we provide the `variable` identifier (implicitly with its resolution context).
// During expansion, the compiler resolves the local `variable` in `main`.
// the macro `print_variable` now compiles - hygiene rules are satisfied!
let variable = "Hello, Macros!";
print_variable!(variable); // Prints "Hello, Macros!"
}
Setting local variable resolution aside for the moment. The other items resolve from the call-site, as mentioned.
The consequence of this is that a macro can work for you as long as the unqualified types, modules, and functions used by the macro are in-scope. To make macros more robust, you may wish to fully qualify such named items in your macro.
// This depends on `HashSet` being in-scope at the call-site.
//
// I find that this is fine to do if the macro is only used in one
// file, but you should prefer the "qualified" route if a macro is
// used in many more places (or is exported from your crate).
macro_rules! make_set_unqualified {
($($elements:expr),*) => {
HashSet::from([$($elements),*])
};
}
// This does not depend on `HashSet` being in-scope at the call-site.
macro_rules! make_set_qualified {
($($elements:expr),*) => {
::std::collections::HashSet::from([$($elements),*])
};
}
fn main() {
// Does not compile, without `use std::collections::HashSet`.
// let set_1 = make_set_unqualified!(1, 2, 3);
// Compiles just fine!
let set_2 = make_set_qualified!(1, 2, 3);
}
There’s a special expansion metavariable for declarative macros -
$crate.You can use this to get a path to the crate that the declarative macro is defined in. So if I wanted to get the path to some local type that the crate the macro was defined in defines, I could use
$crate::CustomType.We will explore exporting, and building resilient macros in more detail in a later article. For now, let’s just focus on how to use declarative macros in isolation.
Now that we understand what macros are, and have a feel for where they execute in the compilation process - let’s look at the syntax used to define and invoke declarative macros.
macro_rules! appears as-if it were a macro, but it has a special syntax that is not replicable. As such, it’s probably best to consider it a built-in language construct.
You can define a new macro using the following syntax.
macro_rules! macro_name_here {
// ... Match Rules Here (at least 1 required) ...
}
This tells the Rust compiler that you want to define a new macro, using the name macro_name_here.
A single “rule” is defined as matcher => transcriber;, where matcher represents the tokens to match with, and transcriber represents the tokens to emit.
The simplest form of this is a macro which matches nothing and expands to nothing.
// This compiles, and you can use it - but it does nothing.
macro_rules! useless_macro {
() => {};
}
fn main() {
useless_macro!();
}
The matcher part may contain literal tokens, and likewise the transcription can be to a pre-defined series of tokens not influenced by the input.
macro_rules! match_literal {
(a) => { println!("Matched a") };
(b) => { println!("Matched b") };
(c) => { println!("Matched c") };
}
fn main() {
match_literal!(a); // "Matched a"
match_literal!(b); // "Matched b"
match_literal!(c); // "Matched c"
// match_literal!(d); // Does not compile!
}
Another step up from this is matching dynamically. To do this, you use what are called “metavariables”.
Metavariables are of the form $name:spec.
nameis allowed to be any identifier or keyword other thancrate; raw identifiers are also permitted.specis the fragment specifier, and it defines the kind of sequence of tokens that this metavariable matches.
The important option here is the fragment specifier. Later in the article, we’ll list the various fragment specifiers, but for now let’s just be aware of a simple one called ident - this matches identifiers and certain keywords (e.g. a, foo, lower_snake, UpperCamel, etc).
macro_rules! match_literal {
(a) => { println!("Matched a") };
(b) => { println!("Matched b") };
(c) => { println!("Matched c") };
($other:ident) => { println!("Matched something else") };
}
fn main() {
match_literal!(a); // "Matched a"
match_literal!(b); // "Matched b"
match_literal!(c); // "Matched c"
match_literal!(d); // "Matched something else"
match_literal!(something_else); // "Matched something else"
match_literal!(fooBarBaz); // "Matched something else"
}
Like with a control-flow
match, the order of the rules matter. If$other:identwere placed at the top, all of these lines would print the same thing, “Matched something else” - that is because it matched that macro rule first.
Additionally, you can use the matched fragments within the transcriber.
macro_rules! match_literal {
(a) => { println!("Matched a") };
(b) => { println!("Matched b") };
(c) => { println!("Matched c") };
($other:ident) => {
// `stringify` converts a matched fragment into a string.
// So if we provided `foo` to this, the string would be "foo".
//
// It works for more than justs metavariables, however.
// See: https://doc.rust-lang.org/stable/std/macro.stringify.html
println!("Matched something else: {}", stringify!($other))
};
}
fn main() {
match_literal!(a); // "Matched a"
match_literal!(b); // "Matched b"
match_literal!(c); // "Matched c"
match_literal!(d); // "Matched something else: d"
}
And of course, you don’t need to match only one thing, you can match multiple things.
macro_rules! convert {
// Match an `ident`, then a literal `as` token, then another `ident`.
//
// `ident` is sufficient for this simple example. Realistically, you may
// want to use other fragment specifiers here.
//
// Let's not get ahead of ourselves, though. :)
($a:ident as $b:ident) => { $b::from($a) };
}
fn main() {
let some_u8: u8 = 10;
let some_u16 = convert!(some_u8 as u16);
println!("u16: {some_u16}");
}
You can make it look kind of like a function by matching literal commas, even.
// This uses the `expr` specifier - "tokens that represent an expression".
//
// Just note that there's nothing special about this syntax. We're literally just
// saying "match an expression, then match a literal comma" and so on - that just happens
// to be the syntax for a regular function call.
macro_rules! like_a_function {
($a:expr, $b:expr, $c:expr) => { println!("{}, {}, {}", $a, $b, $c); };
}
fn function_call() -> String {
String::from("example")
}
fn main() {
let local = true;
like_a_function!(local, 123 + 456, function_call());
}
You may have seen code that uses something other than the () group delimiters at macro invocation sites.
Rust macros don’t allow you to explicitly specify which delimiters are used during invocation. The caller chooses the delimiter, and there’s no requirement that it matches whichever delimiters were used in the matcher rule definition.
That said, for readability we tend to follow some patterns with which we use and where.
| Group Delimiters | Used For… | Example |
|---|---|---|
Parenthesis (()) | Function-call | foo!(a, b, ...) |
Square Brackets ([]) | Sequence Definitions | vec![a, b, ...] |
Curly Brackets ({}) | Impl/Type Definitions | impl_traits!{ a, b, ... } |
Just keep in mind these are for readability only - these are all technically interchangeable.
Well, for the most part. There’s one small difference.
The delimiter used at the invocation-site can affect how the invocation itself is parsed. In particular, “Parenthesis” (()) and “Square Brackets” ([]) generally parse as expressions. But “Curly Brackets” ({}) have special parsing behavior in item and statement contexts (RFC #378).
The major impact of this difference is that curly-bracket-delimited macro invocations don’t require a semicolon at the end when used in these special contexts. This is why we tend to use them for impl blocks and other usually module-scoped items.
macro_rules! impl_stub_display {
($type:ty) => {
impl std::fmt::Display for $type {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
write!(f, "Stubbed Display")
}
}
};
}
// Notice this needs no semicolon at the end?
struct Example;
impl_stub_display! { Example }
// But this one requires the semicolon...
struct OtherExample;
impl_stub_display!(OtherExample);
// So, we tend to stick to {} for these cases.
As you work with macros more and more, you’ll eventually want to perform some more complex matches dealing with repeated elements.
The syntax to use in the matcher for repetitions is $(matcher) separator? operator, which is kind of a mouthful so let’s qualify each of these.
matcheris just another match definition - similar to what you may provide tomacro_rules!.- It can contain literal tokens, metavariables, or nested repetitions.
separatoris optional, and it represents a separator to require between each of thematcherstatements.- This can be any single token, aside from delimiters (
{},[],()) or any of theoperatortokens listed below. - A very common example separator is a comma (
,).
- This can be any single token, aside from delimiters (
operatoris required, and it’s the operator for how many times a match must occur to be valid.?is “zero or one”*is “zero or more”+is “one or more”
Much like with regular matcher rules, you can just match literal tokens.
macro_rules! match_repetition {
( $(a),+ ) => { println!("Matched some a") };
( $(b),+ ) => { println!("Matched some b") };
( $(c),+ ) => { println!("Matched some c") };
}
fn main() {
match_repetition!(a);
match_repetition!(b, b, b);
match_repetition!(c, c);
}
If you would like, you can allow an optional final trailing comma by using another repetition.
macro_rules! match_repetition {
( $(a),+ $(,)? ) => { println!("Matched some a") };
( $(b),+ $(,)? ) => { println!("Matched some b") };
( $(c),+ $(,)? ) => { println!("Matched some c") };
}
fn main() {
match_repetition!(a, );
match_repetition!(b, b, b);
match_repetition!(c, c, );
}
And like with normal macro matches, you can match using metavariables.
macro_rules! match_repetition {
( $(a),+ $(,)? ) => { println!("Matched some a") };
( $(b),+ $(,)? ) => { println!("Matched some b") };
( $(c),+ $(,)? ) => { println!("Matched some c") };
( $($other:ident),+ $(,)? ) => { println!("Matched some other things") };
}
fn main() {
match_repetition!(d, e, f, g, h, i);
match_repetition!(d, e, f, g, h, i, );
}
Which - again, like with normal matches - can be expanded during the transcription phase.
The syntax for transcribing is very similar to matching, $(transcriber) separator? operator.
In the transcriber, there must be at least one metavariable so that the compiler knows which captured data to repeat.
While operator can be any of the repetition operators - which you use makes no difference at the transcription site (the matcher alone is responsible for saying how many matches are needed).
As for the separator, it will appear only between repetitions. It will not appear after the final repetition, or if there are no repetitions.
macro_rules! match_repetition {
( $(a),+ $(,)? ) => { println!("Matched some a") };
( $(b),+ $(,)? ) => { println!("Matched some b") };
( $(c),+ $(,)? ) => { println!("Matched some c") };
( $($other:ident),+ $(,)? ) => {
print!("Matched some other things:");
$(print!(" {}", stringify!($other));)*
println!();
};
}
fn main() {
match_repetition!(d, e, f, g, h, i); // "Matched some other things: d e f g h i"
}
And, if you enjoy pain, your repetitions can even contain other repetitions.
macro_rules! match_repetition {
// An outer repetition separated by commas.
// An inner repetition of whitespace-delimiter `a` identifiers.
( $($(a)+),+ $(,)? ) => {
print!("Matched a bunch of As");
};
}
fn main() {
match_repetition!(a, a a a, a a, a a a a a a a a a);
}
You can find a complete list of fragment specifiers here, but I want to cover the ones I find to be commonly useful (this is not a complete list, but it is most of them).
| Specifier | Brief | Examples |
|---|---|---|
block | A block expression. | {stmt; stmt; expr} |
expr | An expression. | 1 + 2, func(), a, etc. |
ident | An identifier or certain keywords (Disallows _). | lower_snake, UpperCamel, etc. |
item | A function, trait, module, struct, enum, impl, use declaration, etc. | struct Foo {}, enum Bar {} etc. |
lifetime | A lifetime token. | 'a, 'static, etc. |
literal | A literal expression, optionally preceded by -. | 1, true, "string", 'c', -42, etc. |
meta | The contents of a meta attribute, excluding #[]. | derive(Debug), doc = "string", etc. |
pat | A pattern. | Some(x), Struct { foo, bar, .. }, etc. |
path | A type-style path. | foo::bar::baz, relative, ::absolute, etc. |
stmt | A statement. | let foo = expr;, bar();, etc. |
tt | A token tree. | a, 123, { foo(); }, etc. |
ty | A type. | i32, Vec<T>, impl Trait, dyn Trait, etc. |
vis | A visibility qualifier, including an empty match for private visibility. | empty, pub, pub(super), etc. |
In order to avoid parsing ambiguity, the predefined fragment specifiers for declarative macros has some rules for which tokens can follow certain fragments (reference).
For example, an expr fragment cannot be followed by arbitrary tokens, because an expression can end in many ways. As such, there’s a limit of which tokens can match after such a fragment.
// This won't compile - `+` is not allowed to follow after an expression.
//
// We can kind of make sense of this because `+` might be *continuing* an
// expression, so it's possible the greediness of this may be ambiguous, in
// this case. Not so much the end of the expression.
macro_rules! invalid_matcher {
($e:expr + $rest:expr) => { /* ... */ };
}
// However, `2`, `;`, and `=>` *are* allowed to follow an expression.
//
// As such, this is a valid macro and will compile as expected.
macro_rules! valid_matcher {
($e:expr ; $rest:expr) => { /*...*/ };
}
The places you may commonly run into this are:
expr/stmt: Can only be followed by [,,;,=>]ty/path: Can only be followed by [{,[,,,=>,:,=,>,>>,;,|,as,where]
However, if you’re curious about the full set of allowed following tokens, you can check this reference.
Most of these are fairly straightforward, but tt is perhaps a bit unobvious - albeit very useful.
Maybe a simple layman’s definition of a token tree is:
- Either a single token, other than a delimiter (
foo,@,123, etc.), or… - A sequence of tokens using a balanced delimiter group (
(foo @ +),{ a 1 },[ ])
A delimited group counts as one token tree, regardless of how many tokens the group contains. Here’s some code as an example.
macro_rules! match_tt {
( $tt:tt ) => {
println!("Matched a single token tree");
};
}
fn main() {
// Matches a single token.
match_tt!(a);
match_tt!(123);
match_tt!(+);
// Matches a single delimited group.
match_tt!(());
match_tt!({});
match_tt!([]);
// The single delimited group can have any number of tokens (or other token trees even).
match_tt!((a b - +));
match_tt!({123});
match_tt!([ 1 { a (+ - +) }]);
// But note that this isn't a repetition of token trees,
// so an ungrouped list of tokens will *NOT* match.
// match_tt!(123,);
// match_tt!(123 +);
// match_tt!(123 + 123);
}
If you just want to capture and forward arbitrary tokens, you can using repetitions. A common example of this is if you want to write an adapter for something that works like a println! macro.
macro_rules! custom_println {
// We're capturing an arbitrary *sequence* of tokens,
// and then forwarding them to the `println!` macro.
( $($tt:tt)* ) => {
print!("{}:{}:{}: ", file!(), line!(), column!());
println!($($tt)*);
};
}
fn main() {
// This custom println will prefix the file, line, and column.
custom_println!("logging a custom message");
custom_println!("with formatting: {}", 123);
}
There are some pretty powerful patterns surrounding token trees - but with great power comes great responsibility. Don’t do something just because you can, often clever patterns - especially in macros - can lead to very difficult to maintain code.
I tend to mostly only use tt in the manner above - forwarding a repetition of tokens. More complex patterns exist (like recursive TT munchers), but I recommend trying not use or need them. Even if you understand the horrors of your complex macro written today, it’s very likely the future you will not.
That’s all for the basics of writing and using declarative macros - we covered:
- The different kinds of macros available in Rust
- How Rust macros differ from C-style macros
- How to write simple declarative macros
- Several useful supported fragment specifiers
I certainly have more to say about this - for the next article I want to go over some interesting examples, and cover some more subjective things such as how I write maintainable macros.
I’m trying not to write as much as I usually do to keep these articles a bit more approachable. But I would be interested in hearing your feedback! If you have questions or feedback, feel free to reach out on any of my socials listed below, or wherever is most convenient to you.

