Showing posts with label clang_attribute. Show all posts
Showing posts with label clang_attribute. Show all posts

Sep 28, 2026

Where Did My Stack Frame Go? Blocking Tail Calls in Clang and GCC

A crash in g shows a backtrace that jumps straight from main to g, even though main called f and f called g. The debugger is fine. The compiler removed f on purpose.

Tail-call optimization

int g(void);
int f(void) { return g(); }
Calling g is the last thing f does, so f no longer needs its stack frame. At -O2, Clang and GCC compile f to a single jump:
f:
    jmp g
This saves a stack frame and a return. It also removes f from the stack, which breaks: 
  • Backtraces: crash reports and debuggers leave out f. 
  • Profilers: g's time is charged to f's caller. Stack walkers: a logger that skips N frames reports the wrong caller.

Three ways to keep the frame

1. Empty asm volatile after the call (one call site, Clang and GCC)
int f(void) {
  int r = g();
  __asm__ __volatile__("");
  return r;
}
The empty asm produces no instructions. It is volatile, so the compiler can neither delete it nor move it before the call. Because it runs after g() returns, g() is no longer in tail position:
f:
    subq $8, %rsp
    call g
    addq $8, %rsp
    ret
2. __attribute__((disable_tail_calls)) (one function, Clang only)
__attribute__((disable_tail_calls))
int f(void) { return g(); }
This states the intent clearly and covers every call inside f.

3. -fno-optimize-sibling-calls (whole file, Clang and GCC) This flag fits profiling builds. For a single function, it turns off too much. Inlining still removes frames
If the compiler inlines f into its caller, f has no frame to keep. When the frame must exist, also add __attribute__((noinline)).
 

Cost

Each blocked tail call costs a call/ret pair and one stack frame instead of a jmp. This matters only on very hot paths, or in deep recursion that relied on tail calls to avoid overflowing the stack.

Demo

Paste this into Compiler Explorer (godbolt.org) and compile with -O2:
int g(void);
int with_tail_call(void)    { return g(); }
int without_tail_call(void) { int r = g(); __asm__ __volatile__(""); return r; }
with_tail_call compiles to jmp g. without_tail_call compiles to call g followed by ret. 
Verified with Clang 21 and GCC 15 on x86-64.

Takeaway

  • One call site: add __asm__ __volatile__("") after the call.
  • One function (Clang): use __attribute__((disable_tail_calls)).
  • A whole file: compile with -fno-optimize-sibling-calls.
  • To guarantee the frame: also add noinline.

Dec 27, 2025

[C++] coroutine cheat sheet - 3

Reference:
[C++] Object Lifetimes reading minute



When HALO (Heap Allocation Elision Optimization) could happen.

  1. The Lifetime is Strictly Nested
    The compiler must be able to prove that the coroutine's lifetime ends before the caller's execution finishes.
    If the coroutine object (the "handle" or "task") is returned and its destruction point cannot be determined at compile-time, the compiler must play it safe and use the heap.
  2. The Coroutine State Size is Known
    The compiler needs to know exactly how much space the coroutine requires (including captured variables and promise objects) at the call site.
    This usually requires the coroutine body to be visible to the compiler
    (i.e., in the same translation unit or available via Link Time Optimization).
  3. Use of std::get_return_object_on_allocation_failure (Optional but relevant)
    If our promise_type defines this static member function, it signals to the compiler how to handle allocation failures.
    While this doesn't force an opt-out, it changes the allocation strategy to be more robust.

The HALO Optimization Process:

The optimization works roughly like this:

  • Analysis: The compiler looks at the co_await and destruction points of the coroutine object.
  • Inlining: It attempts to inline the coroutine logic into the caller.
  • Elision: If the compiler sees that the coroutine state does not "escape" the function, it replaces operator new with a local stack allocation.

How to Encourage the Compiler to Opt-Out

Since we cannot explicitly keywords like noheap, we have to "help" the compiler's optimizer:
  • Keep the coroutine local: Avoid passing the coroutine handle to other threads or storing it in global containers.
  • Enable High Optimization: HALO typically requires -O2 or -O3 (GCC/Clang) or /O2 (MSVC).
  • Inline the Coroutine: Define the coroutine in a header or the same file where it is called so the compiler can see the full lifecycle.
Task my_coroutine() {
    co_return 42;
}

void caller() {
    auto t = my_coroutine(); // Compiler can see 't' lives only here
    // ... do something ...
} // 't' is destroyed here; Heap allocation likely elided. 


Limitations

  • Dynamic Dispatch: If we call a coroutine through a virtual function or a function pointer, the compiler usually cannot perform HALO.
  • Tail Calls: Complex chains of coroutines can sometimes make it difficult for the compiler to prove nested lifetimes.

The "Manual" Opt-Out (Custom Allocators)

If we cannot rely on the compiler's optimization (e.g., in embedded systems), 
we can manually opt-out of the default heap by overloading operator new in our promise_type.

Using a Static Buffer or Arena:

We can provide a custom operator new that pulls from a pre-allocated memory pool or a stack-based arena, effectively bypassing the system heap.
struct promise_type {
    // Overloading new allows we to use a custom allocator
    void* operator new(std::size_t size) {
        return my_custom_arena.allocate(size);
    }
    void operator delete(void* ptr) {
        my_custom_arena.deallocate(ptr);
    }
    // ... other promise members
};


How to Verify if Elision Happened

Since HALO is an optimization, it can be fragile. We can verify it using these tricks:
  • Print in Custom New: Add a printf inside our promise_type::operator new. If it doesn't print during execution at -O3, the compiler successfully elided the call.
  • Compiler Explorer (Assembly): Check for the absence of call operator new or malloc in the generated assembly.
  • Clang-Specific Attributes: Clang is experimenting with attributes like [[clang::coro_inplace_task]] to make this elision more deterministic, though this is not standard C++20. (reference: Language Extension for better, more deterministic HALO for C++ Coroutines)
#include <coroutine>
#include <iostream>
#include <array>

struct StaticTask {
    struct promise_type {
        // 1. Intercept the arguments of the coroutine function
        // This allows 'operator new' to see the buffer passed to the coroutine
        void* operator new(std::size_t size, std::span<std::byte> buffer) {
            if (size > buffer.size()) {
                throw std::bad_alloc();
            }
            std::cout << "Allocating " << size << " bytes from stack buffer\n";
            return buffer.data();
        }

        // Must provide a matching delete (even if it does nothing)
        void operator delete(void*, std::size_t) {}

        StaticTask get_return_object() { return {std::coroutine_handle<promise_type>::from_promise(*this)}; }
        std::initial_suspend initial_suspend() { return {}; }
        std::final_suspend final_suspend() noexcept { return {}; }
        void return_void() {}
        void unhandled_exception() { std::terminate(); }
    };

    std::coroutine_handle<promise_type> handle;
};

// Usage
StaticTask my_coro(std::span<std::byte> storage) {
    std::cout << "Coroutine running!\n";
    co_return;
}

int main() {
    std::array<std::byte, 1024> stack_space; 
    auto task = my_coro(stack_space); // Passes the buffer to 'operator new'
}

In C++20, the operator new for a coroutine is uniquely powerful because the compiler performs a special lookup.
It doesn't just look for a standard void* operator new(size_t); it looks for an overload that matches the entire signature of our coroutine function.

Why it takes the buffer as an argument
When we call a coroutine like my_coro(some_buffer), the compiler needs to allocate space for the "coroutine frame" 
(which holds local variables and the promise).

To give us total control, the C++ standard says:

The compiler will first try to find an operator new in our promise_type that takes (std::size_t, Args...),
where Args... are the exact types passed to the coroutine function.

If it finds this "matching" version, it calls it and passes the arguments we provided in the function call.

This is the "magic hook" that allows us to pass a specific memory source (like a stack-based span or a custom Arena&) directly into the allocation logic.

The Lookup Mechanics

The compiler follows this priority list when it sees a coroutine call:

Priority   Signature the compiler looks for Description
  1. (Best)  operator new(size_t, P1, P2...)
    Takes the size plus all coroutine arguments (P1,P2).
  2. operator new(size_t)                   
    The standard class-specific allocator.
  3. ::operator new(size_t)                 
    Global heap allocation. (Fallback)
Note: For member functions, the first argument after size_t is actually the this pointer, followed by the function arguments.



struct promise_type {
    std::span<std::byte> m_buffer;

    // 1. Used to POSITION the memory (Allocation)
    void* operator new(std::size_t size, std::span<std::byte> buffer) {
        if (size > buffer.size()) throw std::bad_alloc();
        return buffer.data();
    }

    // 2. Used to INITIALIZE the promise (Construction)
    // The compiler sees that the coroutine was called with a span,
    // so it looks for a constructor that accepts it.
    promise_type(std::span<std::byte> buffer) : m_buffer(buffer) {
        std::cout << "Promise initialized with buffer of size " << m_buffer.size() << "\n";
    }

    // ... rest of promise_type members ...
};

[[clang::coro_await_elidable_argument]] 

attribute is a specialized Clang-specific optimization hint designed to reduce the overhead of asynchronous programming. It is primarily used to enable HALO (Heap Allocation Elision Optimization) for coroutines.

This attribute is applied to a parameter of a function (usually an operator await). It tells the compiler that the coroutine being passed as an argument is a temporary that will not outlive the current function. This gives the compiler a "green light" to:
  • Elide the heap allocation: Instead of putting the coroutine frame on the heap, it puts it on the caller's stack.
  • Inline the coroutine: It allows for better devirtualization and inlining of the coroutine's lifecycle.

Sep 2, 2025

[clang] attribute enable_if

enable-if
int isdigit(int c);
int isdigit(int c) __attribute__((enable_if(c <= -1 || c > 255, "chosen when 'c' is out of range"))) __attribute__((
	unavailable("'c' must have the value of an unsigned char or EOF")));

void foo(char c) {
  isdigit(c);
  isdigit(10);
  isdigit(-10);  // results in a compile-time error; as like static_assert.
}

Jul 13, 2025

[C++][Rust] default Lifetime annotation from Rust

  1. The first rule is that the compiler assigns a lifetime parameter to each parameter that’s a reference. In other words, a function with one parameter gets one lifetime parameter: fn foo<'a>(x: &'a i32); a function with two parameters gets two separate lifetime parameters: fn foo<'a, 'b>(x: &'a i32, y: &'b i32); and so on.
  2. The second rule is that, if there is exactly one input lifetime parameter, that lifetime is assigned to all output lifetime parameters: fn foo<'a>(x: &'a i32) -> &'a i32.
  3. The third rule is that, if there are multiple input lifetime parameters, but one of them is &self or &mut self because this is a method, the lifetime of self is assigned to all output lifetime parameters. This third rule makes methods much nicer to read and write because fewer symbols are necessary.

Consider the idea and apply to C++ with 
gnu::lifetimebound
#include <map>
#include <string>

using namespace std::literals;

// Returns m[key] if key is present, or default_value if not.
template<typename T, typename U>
const U &get_or_default(const std::map<T, U> &m [[clang::lifetimebound]],
                        const T &key, /* note, not lifetimebound */
                        const U &default_value [[clang::lifetimebound]]) {
  if (auto iter = m.find(key); iter != m.end()) return iter->second;
  else return default_value;
}

int main() {
  std::map<std::string, std::string> m;
  // warning: temporary bound to local reference 'val1' will be destroyed
  // at the end of the full-expression
  const std::string &val1 = get_or_default(m, "foo"s, "bar"s);

  // No warning in this case.
  std::string def_val = "bar"s;
  const std::string &val2 = get_or_default(m, "foo"s, def_val);

  return 0;
} 
Output:
<source>:19:55: warning: temporary bound to local reference 'val1' will be destroyed at the end of the full-expression [-Wdangling]
   19 |   const std::string &val1 = get_or_default(m, "foo"s, "bar"s);

Mar 6, 2024

[C++] avoid tail call based on different compiler


// BLOCK_TAIL_CALL_OPTIMIZATION
//
// Instructs the compiler to avoid optimizing tail-call recursion. This macro is
// useful when you wish to preserve the existing function order within a stack
// trace for logging, debugging, or profiling purposes.
//
// Example:
//
//   int f() {
//     int result = g();
//     BLOCK_TAIL_CALL_OPTIMIZATION();
//     return result;
//   }
#if defined(__pnacl__)
#define BLOCK_TAIL_CALL_OPTIMIZATION() if (volatile int x = 0) { (void)x; }
#elif defined(__clang__)
// Clang will not tail call given inline volatile assembly.
#define BLOCK_TAIL_CALL_OPTIMIZATION() __asm__ __volatile__("")
#elif defined(__GNUC__)
// GCC will not tail call given inline volatile assembly.
#define BLOCK_TAIL_CALL_OPTIMIZATION() __asm__ __volatile__("")
#elif defined(_MSC_VER)
#include 
// The __nop() intrinsic blocks the optimisation.
#define ABSL_BLOCK_TAIL_CALL_OPTIMIZATION() __nop()
#else
#define ABSL_BLOCK_TAIL_CALL_OPTIMIZATION() \
  if (volatile int x = 0) {                 \
    (void)x;                                \
  }
#endif

Jan 2, 2023

[C++/Rust] pragma nounroll

Reference:
https://stackoverflow.com/questions/74979866/how-to-stop-clang-from-overexpanding
https://clang.llvm.org/docs/AttributeReference.html#pragma-unroll-pragma-nounroll

Use #pragma nounroll to control loop unrolling optimization; for Rust counter part:

https://docs.rs/unroll/latest/unroll/

#include <iostream>
typedef long xint;
template<int N>
struct foz {
    template<int i=0>
    static void foo(xint t) {
        // #pragma nounroll
        for (int j=0; j<10; ++j) {
            foo<i+1> (t+j);
        }
    }
};

template<>
template<>
void foz<4>::foo<4>(xint t) {
    std::cout << t;
}

int main() {
    foz<4>::foo<0>(0);
}

Sep 9, 2021

[C++] Safer Usage Of C++ note

Reference:
Safer Usage Of C++

CLang user manual:

https://clang.llvm.org/docs/UsersManual.html

https://clang.llvm.org/docs/ClangCommandLineReference.html


Enable flags:

-fno-exceptions
-ftrapv
-fwrapv
fsanitize=signed-integer-overflow
-Wdangling-gsl

-fno-delete-null-pointer-checks (named as such for historical reasons) that defines null pointer dereferences. With this flag, dereferences of null are never optimized away.


MiraclePtr:

https://youtu.be/ohlxw5kDn-k

https://docs.google.com/presentation/d/1QvfZXx5HdUl0IdkBcrx-NM0ua-PVcTi2jNx0Sf-n8Fo/edit#slide=id.gab22a695b8_0_1


scpptool 

is a command line tool to help enforce a memory and data race safe subset of C++. 

https://github.com/duneroadrunner/scpptool


"SaferCPlusPlus" is essentially a collection of safe data types intended to facilitate memory and data race safe C++ programming.

https://github.com/duneroadrunner/SaferCPlusPlus

https://github.com/duneroadrunner/SaferCPlusPlus-AutoTranslation2


StarScan

Heap scanning use-after-free prevention

https://source.chromium.org/chromium/chromium/src/+/master:base/allocator/partition_allocator/starscan/README.md


MiraclePtr aka raw_ptr aka BackupRefPtr

https://chromium.googlesource.com/chromium/src/+/ddc017f9569973a731a574be4199d8400616f5a5/base/memory/raw_ptr.md


Pointer Safety Ideas

https://docs.google.com/document/d/1qsPh8Bcrma7S-5fobbCkBkXWaAijXOnorEqvIIGKzc0/edit#


P1705R1

Enumerating Core Undefined Behavior

http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2019/p1705r1.html


Automatic Reference Counting

https://en.wikipedia.org/wiki/Automatic_Reference_Counting


Blink GC API reference

https://chromium.googlesource.com/chromium/src/+/refs/heads/main/third_party/blink/renderer/platform/heap/BlinkGCAPIReference.md

https://docs.google.com/presentation/d/1XPu03ymz8W295mCftEC9KshH9Icxfq81YwIJQzQrvxo/edit#slide=id.p


2 basic types of memory safety

spatial:

The program will behave in a defined and safe way if it accesses memory outside valid bounds.

Examples include array bounds, struct and union field access, and iterator access.


temporal:

The program will behave in a defined and safe way if it accesses memory when that memory is not valid at the time of the access.

Examples include use after free (UAF), double-free, use before initialization, and use after move (UAM).


[[clang::lifetimebound]] 

https://clang.llvm.org/docs/AttributeReference.html#lifetimebound


ABSL

Use absl::variant Instead Of enums for state machines


Jan 26, 2021

[C++] trivial_abi and borrow ownership

An interesting concept comes up from Rust's borrow ownership.

In C++, the unique_ptr can only be moved. While instead of using heavy-weighted shared_ptr
(i.e 2 words size, heap allocated control block, atomic ref-counting, weak_ptr etc.), the only way of borrow the ownership from unique_ptr is by passing raw pointer from unique_ptr.

Here's the tweet from Prof. John Regehr talking about this:
https://twitter.com/johnregehr/status/1351333748277592065
Follow up from s.o:
https://stackoverflow.com/a/30013891
i.e in C++, the borrow concept can't be made in compile time, but in runtime, which unlike Rust.

Another to mention about is trivial_abi.
AMD64 ABI for C++: https://www.uclibc.org/docs/psABI-x86_64.pdf

While call by value,

Quote from ABI:
If a C++ object has either a non-trivial copy constructor or a non-trivial destructor, it is passed by invisible reference (the object is replaced in the parameter list by a pointer […]).

And in Itanium C++ ABI document, quote:

non-trivial for the purposes of calls
A type is considered non-trivial for the purposes of calls if: it has a non-trivial copy constructor, move constructor, or destructor, or all of its copy and move constructors are deleted.
This definition, as applied to class types, is intended to be the complement of the definition in [class.temporary]p3 of types for which an extra temporary is allowed when passing or returning a type. A type which is trivial for the purposes of the ABI will be passed and returned according to the rules of the base C ABI, e.g. in registers; often this has the effect of performing a trivial copy of the type.

That is call by value with non-trivial type introduce double indirection during the call while the type could potentially not fit into the register for callee.  By double indirection meaning the argument is reference to the temporary r-value created on caller's stack.
i.e potential virtual pointer to v-table increases the size of type.


Reference: https://www.raywenderlich.com/615-assembly-register-calling-convention-tutorial

Assembly code can be found in godbolt:
https://godbolt.org/z/s1TPea6sx


#include <cstdio>

#define TRIVIAL_ABI __attribute__((trivial_abi))

template <class T> T incr(T obj) {
  obj.value += 1;
  puts("before exit incr func");
  return obj;
}

struct Up1 {
  int value;
  Up1() = default;
  Up1(const Up1& u) : value(u.value) { puts("Up1 copy constructor"); }
  ~Up1() { printf("detroyed Up1 value: %d\n", value); }
};

struct TRIVIAL_ABI Up2 {
  int value;
  Up2() = default;
  Up2(const Up1& u) : value(u.value) { puts("Up2 copy constructor"); }
  ~Up2() { printf("detroyed Up2 value: %d\n", value); }
};

template Up1 incr(Up1);
template Up2 incr(Up2);

auto main() -> int {
  auto u1 = Up1{};
  puts("before call incr func for u1");
  incr(u1);

  printf("\n\n");

  auto u2 = Up2{};
  puts("before call incr func for u2");
  incr(u2);
}

Output:
before call incr func for u1
Up1 copy constructor; temporary object creatd;
before exit incr func
Up1 copy constructor; temporary object creatd;
detroyed Up1 value: 1
detroyed Up1 value: 1


before call incr func for u2
// No temporary object created, all in register.
before exit incr func
detroyed Up2 value: 1
detroyed Up2 value: 1
detroyed Up2 value: 0
detroyed Up1 value: 0

Jul 11, 2019

[clang] Catching use-after-move bugs with Clang's consumed annotations

Reference:
https://awesomekling.github.io/Catching-use-after-move-bugs-with-Clang-consumed-annotations/

Clang9 Attributes:
https://clang.llvm.org/docs/AttributeReference.html#consumed-annotation-checking

The reason once we've defined destructor for our type compiler will not generate default move constructor/move assignment which ask the designer to explicit define those two member functions due to move semantic,  i.e what should the r-value object's data members being handled? If the data member is the resource owner , it should be nullptred(by move constructor/move assignment), which allows the type's destructor delete the nullptr(i.e a no-op).
However, if any member functions which deref the nullptred data member will crash the process at run-time.

Can this being caught at compile time in C++, like Rust did?
Clang provides “Consumed Annotation Checking”

code example:
class [[clang::consumable(unconsumed)]] CleverObject {
public:
    CleverObject() {}
    CleverObject(CleverObject&& other) { other.invalidate(); }

    [[clang::callable_when(unconsumed)]]
    void do_something() { assert(m_valid); }

private:
    [[clang::set_typestate(consumed)]]
    void invalidate() { m_valid = false; }

    bool m_valid { true };
};

int main(int, char**)
{
    CleverObject object;
    auto other = std::move(object);
    object.do_something();
    return 0;
}
Realworld Usage:
https://github.com/SerenityOS/serenity/blob/master/AK/NonnullRefPtr.h