The strict aliasing situation is pretty bad (2016) (opens in new tab)

(blog.regehr.org)

74 pointspkkm1y ago58 comments

58 comments

36 comments · 6 top-level

jandrewrogers1y ago· 9 in thread

This is an issue that is ignored by just about everyone in practice. The reality is that most developers have subconsciously internalized the compiler behavior and assume that will always hold. And they are mostly right, I’ve only seen a few cases where this has caused a bug in real systems over my entire career. I try, to the extent possible, to always satisfy the requirements of strict aliasing when writing code. It is difficult to determine if I’ve been successful in this endeavor.

Here is why I don’t blame the developers: writing fast, efficient systems code that satisfies the requirements of strict aliasing as defined by C/C++ is surprisingly difficult. It has taken me years to figure out the technically correct incantations for every weird edge case such that they always satisfy the requirements of strict aliasing. The code gymnastics in some cases are entirely unreasonable. In fairness, recent versions of C++ have been adding ways to express each of these cases directly, eliminating the need to use obtuse incantations. But we still have huge old code bases that assume compiler behavior, as was the practice for decades.

I am not here to attribute blame, I think it the causes are pretty diffuse honestly. This is just a part of the systems world we failed to do well, and it impacts the code we write every day. I see strict aliasing violations in almost every code base I look at.

spacechild11y ago

> In fairness, recent versions of C++ have been adding ways to express each of these cases directly, eliminating the need to use obtuse incantations.

In particular, C++20 gave us std::bit_cast (https://en.cppreference.com/w/cpp/numeric/bit_cast) for type punning and C++23 added std::start_life_time_as (https://en.cppreference.com/w/cpp/memory/start_lifetime_as) for interpreting raw bytes as an object.

helltone1y ago

std::start_life_time_as and std::launder honestly feel like a massive hack

0823498723498721y ago

One of the issues that worked against Euclid's adoption was that its compiler strictly disallowed aliasing. That said, https://dl.acm.org/doi/pdf/10.5555/800078.802513 claims that while they tended to write potentially-aliased code at first, after one made the Euclid compiler happy, subsequent development wasn't likely to reintroduce it.

https://en.wikipedia.org/wiki/Euclid_(programming_language)

tialaramex1y ago

The insight in languages like Rust is that aliasing is actually fine if we can guarantee all the aliases are immutable and that's facilitated by default reference immutability. [These alias related] Bugs only arise when you have mutable aliasing which is why that doesn't exist in safe Rust.

That paper also highlights that checking is crucial, their initial Euclid compiler just required that there's no aliasing, but never checked. So of course programmers will make mistakes and without the checks those mistakes leak into running code. The finished compiler checked, which means the mistake won't even compile.

Shifting left in this way is huge, WUFFS shifts bounds misses left - when you write code which can have a bounds miss in C of course it just does have a bounds miss at runtime, there's a stray read or overwrite and chaos results maybe it's Remote Code Execution, in Rust the miss panics at runtime - maybe a Denial of Service or at least a major inconvenience. But in WUFFS it won't compile - you find out about your bug likely before it gets sent out for code review.

Most software can't be written in WUFFS, but "most" is doing a lot of work there, plenty of code which should be in WUFFS or an analogous language is not, meaning mistakes are not shifted left.

gpderetta1y ago

Indeed, another problem is that we have no tools, other than very imperfect linters/compiler warnings, to identify aliasing violations. Even today I don't think sanitizers can catch most cases.

RossBencina1y ago

More often than not when I realise that I am violating strict aliasing it is because I am doing something that I want to do and the language is not going to let me. Much hand wringing, language lawyering and time wasting typically follows.

1 more reply

Arch-TK1y ago

> The reality is that most developers have subconsciously internalized the compiler behavior and assume that will always hold.

I blame this on how people like to teach C and present C.

It's very important that the second anyone conceives of the idea of learning C that they first off informed that trying things and seeing what happens is a highly unreliable method of learning how C programs behave and that C is not a high level assembly language.

If you teach C in relation to the abstract machine instead of any real world machine you will understandably scare off most people. Which is good, since most people shouldn't be learning or writing C. It's a language which can barely be written correctly even by people with the necessary self discipline to only write code they're 100% certain is well defined.

> It is difficult to determine if I’ve been successful in this endeavor.

Why is your program so full of casts between pointer types that you have difficulty determining if you've avoided strict aliasing?

Yes, if you treat C as a high level assembly language (like the linux kernel likes to do) then it becomes difficult to reason about the behaviour of your programs where 50% of them are in the grey area of uncertainty of whether they're well defined or not.

If you are forced to write C in a non-learning context, don't write any line of code unless you're certain you could tell someone which parts of the standard describe its behaviour.

> Here is why I don’t blame the developers: writing fast, efficient systems code that satisfies the requirements of strict aliasing as defined by C/C++ is surprisingly difficult.

C/C++ isn't a language. So I will stick to C because I don't know nor care about C++.

That being said, it's not hard to write efficient C which satisfies the requirements of strict aliasing except when you're dealing with idiotic APIs like bind or connect. Most code by default, assuming you use appropriate algorithms and data structures, is performant. The only time it becomes difficult with regards to strict aliasing is if you're micro optimizing.

While non-trivial, the case of converting between unsigned long and float shown in the article is entirely possible to do with completely safe C constructs. Likewise serialization/deserialization of binary data never requires coming close to aliasing unless you're dealing with a "native" endian protocol. In the case of general serialisation and deserialisation, compilers will reliably optimise such operations into one or two instructions (depending on whether you're decoding same-endianness or not).

jandrewrogers1y ago

> Why is your program so full of casts between pointer types that you have difficulty determining if you've avoided strict aliasing?

I write database storage engines. Most of the runtime address space is being dynamically paged to storage directly by user space. You can't use mmap() for this. Consequently, objects don't have a fixed address over their lifetime and what a pointer actually points to is not always knowable at compile-time. These are all things that have to be dynamically resolved at runtime with zero copies in every context the memory might be touched. Fairly standard high-performance database stuff. The intrinsic ambiguity about the contents of a memory address create many opportunities to inadvertently create strict aliasing violations.

I've been doing it a long time, so I know the correct incantation for virtually every difficult strict aliasing edge case. Most developers are ignorant of at least some of these incantations because they are surprisingly difficult to lookup, it took me years to figure out some of them. When developers don't know they tend to YOLO it and hope the compiler does the desired thing. Which mostly works in practice, until it doesn't.

Recent versions of C++ have added explicit helper functions, which is a big improvement. Most developers don't know the code incantation required to reliably achieve the same effect as std::start_lifetime_as and they shouldn't have to.

2 more replies

gpderetta1y ago

Well, it is more complicated than that.

First of all compilers disagree on many interpretations and consequences of abstract machine rules. Also compilers have bugs.

So a proficient C/C++ programmer does have to learn what compilers actually do in practice and what they guarantee beyond the standard (or how they differ from it).

> C/C++ isn't a language.

It isn't, but it is a family of languages that share a lot of syntax and semantics.

2 more replies

Gabriel541y ago· 8 in thread

Forgive me my ignorance, but if I write

  int foo(int *x) {
    *x = 0;
    // wait until another thread writes to *x
    return *x;
  }

Can the C compiler really optimize foo to always return 0? That seems extremely unintuitive to me.

fweimer1y ago

How do you accomplish the waiting operation? If it does not synchronize with the other thread, the compiler will optimize away the load. This isn't too surprising once you assume that not every *x in the source code will result in a memory access instruction. I would even say that most C programmers expect such basic optimizations to happen, although they might not always like the consequences.

lmm1y ago

> Can the C compiler really optimize foo to always return 0?

Yes

> That seems extremely unintuitive to me.

C compilers are extremely unintuitive. This is a relatively sane case, they do things that are much more surprising than this.

badmintonbaseba1y ago

In a multithreaded, hosted userspace program the wait operations should synchronize with another thread. This involves inserting optimization barriers that are understood by the compiler, therefore it can't optimize the this case to always return 0.

AlexandrB1y ago

In embedded this situation is quite common when x points to a hardware register. The typical solution is to declare x as volatile[1] which tells the compiler to omit these optimizations.

It's very common for beginner embedded programmers to forget to do this and spend hours debugging why the register doesn't change when it should.

[1] https://en.m.wikipedia.org/wiki/Volatile_(computer_programmi...

ajross1y ago

No, because in practice that "wait until" operation will act as a memory barrier. The obvious one is a function call. Functions are allowed to have side effects, one possible side effect is to change the value pointed to by an externally-received pointer.

At lower levels, you might have something like an IPC primitive there, which would be protected by a spinlock or similar abstraction, the inline assembly for which will include a memory barrier.

And even farther down still, the memory pointed to by "x" might be shared with another async context entirely and the "wait for" operation might be a delay loop waiting on external hardware to complete. In that case this code would be buggy and you should have declared the data volatile.

sapiogram1y ago

> No, because in practice that "wait until" operation will act as a memory barrier.

This is a wrong, a memory barrier would not salvage this code from UB. The read from `x` must at the very least be synchronized, and there might be other UB lurking as well.

1 more reply

marcosdumay1y ago

That's the most straight-forward example of undefined behavior badness you'll find. Things on practice are usually way less intuitive than this (mostly because people notice and avoid writing those straight-forward problems).

tempodox1y ago

Yes. You'd have to use

  int volatile *x

as the parameter to get the changes from a different thread.

sapiogram1y ago· 6 in thread

Has anyone measured the performance impact of the -fno-strict-aliasing flag? How much real-world performance are we really gaining from all this mess?

sestep1y ago

Not sure if this is exactly the same scope as what you're asking about, but here's an ESSE '21 paper titled "The Impact of Undefined Behavior on Compiler Optimization": https://doi.org/10.1145/3501774.3501781

sapiogram1y ago

It's paywalled, is there a way to actually read it?

gpderetta1y ago

Depends a lot on the application. For many it matters little, but for some (mostly numerical), it can matter a lot.

twoodfin1y ago

Obviously it’s going to vary from program to program. And you always have to be skeptical that removing the safety for performance hasn’t given you a faster but faulty program.

That being said, my intuition matches what little anecdotal data I’ve seen from real perf-sensitive systems, and I’d ballpark 10-15% where it matters.

lmm1y ago

Real-world performance: not enough to be measurable, certainly remotely enough to make up for the time we lose to debugging.

But no-one cares about real-world performance, people pick C and pick a C compiler because they want the thing that's fastest on artificial microbenchmarks.

73kl4453dz1y ago

Lack of aliasing was historically fortran's advantage over C.

RossBencina1y ago· 5 in thread

Interesting that the article doesn't even entertain the obvious solution: remove strict aliasing requirements from the standards.

throw161803391y ago

There are no obvious changes to a widely used programming language standard.

Even small changes often require years and many revisions to be accepted - burnout is common. You would need to build a consensus that this change is desirable - that's highly unlikely at best. Strict aliasing has been widely implemented since the 1990s and many compilers benefit from the rules; many compiler vendors are on the committee. You'd have to convince them that they should make their customer's code slower.

What might be achievable, however, is some kind of technical report on undefined or implementation defined behavior. Many compilers have options that allow programs with some undefined behavior to behave as the user would expect. Microsoft's C and C++ compilers, for example, don't enforce strict aliasing and allow some forms of integer overflow in loop conditionals. There would be substantial value in defining a common profile for these options. It would still be an uphill battle to get it through the committee, though.

gpderetta1y ago

A while ago the C++ committee tried to standardize function argument evaluation order. It actually made it to the draft standard, but it had to be reverted when it was presented with real world performance regressions.

If we can't even get that, I doubt strict aliasing will ever be voted out.

bluGill1y ago

Compiler writers tell me that it makes a big difference to optimization. I am careful to never cast anything in ways that there are problems and so I run with strict aliasing. My project started in 2010 though, so we had plenty of prior best practices to help us know better and no legacy code that is hard to refactor to make correct. We have had out share of memory issues, but never anything that could be blamed on strict aliasing.

zokier1y ago

on the other hand, it does say

> If I were writing correctness-oriented C that relied on these casts I wouldn’t even consider building it without -fno-strict-aliasing.

Arch-TK1y ago

"correctness-oriented C" definitionally cannot consider "[relying] on those casts".

icedchai1y ago· 2 in thread

I learned C on the Amiga, back in the late 80's, and the OS made heavy use of "OO-ish" physical subtyping with structs everywhere. I don't think anybody even thought about strict aliasing violations.

ajross1y ago

Compilers in the 1980's really weren't sophisticated enough to have this problem. A function call was a hard barrier that was going to spill all GPRs, inlining was almost unheard of. What the code did was what you saw, and if you had an aliased pointer it's because that's what you wanted.

And when it became an issue c. late 90's, it was actually "NO strict aliasing" that was the point of contention. Optimizers were suddenly able to do all sorts of magic, and compiler authors realized they were getting tripped up by the inability (c.f. the halting problem) to know for sure that this arbitrary pointer wasn't scribbling over the memory contents they were trying to optimize. You'd get better (often much better) code with -fno-strict-aliasing, which was tempting enough to turn it on and hope for better analysis tools to come along and save us from the resulting bugs.

We're still waiting, alas.

flohofwoe1y ago

The Amiga C compilers most likely didn't do a lot of optimizations where strict aliasing would matter though (at least from what I remember it was pretty straight forward, a memory read or write in C typically resulted in a memory read or write in assembly).

Basically, C code compiled to assembly in the Amiga era looked much more straightforward than the output produced by modern C compilers (with optimizations enabled at least), you could put both side by side and see a near 1:1 relationship between the C code and the assembly code (maybe also because the Motorola 68000 seems to have taken a lot of inspiration from the PDP instruction set).

dang1y ago

Discussed at the time:

The Strict Aliasing Situation Is Pretty Bad - https://news.ycombinator.com/item?id=11288665 - March 2016 (67 comments)

j / k navigate · click thread line to collapse

58 comments

36 comments · 6 top-level

jandrewrogers1y ago· 9 in thread

spacechild11y ago

> In fairness, recent versions of C++ have been adding ways to express each of these cases directly, eliminating the need to use obtuse incantations.

helltone1y ago

std::start_life_time_as and std::launder honestly feel like a massive hack

0823498723498721y ago

https://en.wikipedia.org/wiki/Euclid_(programming_language)

tialaramex1y ago

Most software can't be written in WUFFS, but "most" is doing a lot of work there, plenty of code which should be in WUFFS or an analogous language is not, meaning mistakes are not shifted left.

gpderetta1y ago

Indeed, another problem is that we have no tools, other than very imperfect linters/compiler warnings, to identify aliasing violations. Even today I don't think sanitizers can catch most cases.

RossBencina1y ago

1 more reply

Arch-TK1y ago

> The reality is that most developers have subconsciously internalized the compiler behavior and assume that will always hold.

I blame this on how people like to teach C and present C.

> It is difficult to determine if I’ve been successful in this endeavor.

Why is your program so full of casts between pointer types that you have difficulty determining if you've avoided strict aliasing?

If you are forced to write C in a non-learning context, don't write any line of code unless you're certain you could tell someone which parts of the standard describe its behaviour.

> Here is why I don’t blame the developers: writing fast, efficient systems code that satisfies the requirements of strict aliasing as defined by C/C++ is surprisingly difficult.

C/C++ isn't a language. So I will stick to C because I don't know nor care about C++.

jandrewrogers1y ago

> Why is your program so full of casts between pointer types that you have difficulty determining if you've avoided strict aliasing?

2 more replies

gpderetta1y ago

Well, it is more complicated than that.

First of all compilers disagree on many interpretations and consequences of abstract machine rules. Also compilers have bugs.

So a proficient C/C++ programmer does have to learn what compilers actually do in practice and what they guarantee beyond the standard (or how they differ from it).

> C/C++ isn't a language.

It isn't, but it is a family of languages that share a lot of syntax and semantics.

2 more replies

Gabriel541y ago· 8 in thread

Forgive me my ignorance, but if I write

  int foo(int *x) {
    *x = 0;
    // wait until another thread writes to *x
    return *x;
  }

Can the C compiler really optimize foo to always return 0? That seems extremely unintuitive to me.

fweimer1y ago

lmm1y ago

> Can the C compiler really optimize foo to always return 0?

Yes

> That seems extremely unintuitive to me.

C compilers are extremely unintuitive. This is a relatively sane case, they do things that are much more surprising than this.

badmintonbaseba1y ago

AlexandrB1y ago

In embedded this situation is quite common when x points to a hardware register. The typical solution is to declare x as volatile[1] which tells the compiler to omit these optimizations.

It's very common for beginner embedded programmers to forget to do this and spend hours debugging why the register doesn't change when it should.

[1] https://en.m.wikipedia.org/wiki/Volatile_(computer_programmi...

ajross1y ago

At lower levels, you might have something like an IPC primitive there, which would be protected by a spinlock or similar abstraction, the inline assembly for which will include a memory barrier.

sapiogram1y ago

> No, because in practice that "wait until" operation will act as a memory barrier.

This is a wrong, a memory barrier would not salvage this code from UB. The read from `x` must at the very least be synchronized, and there might be other UB lurking as well.

1 more reply

marcosdumay1y ago

tempodox1y ago

Yes. You'd have to use

  int volatile *x

as the parameter to get the changes from a different thread.

sapiogram1y ago· 6 in thread

Has anyone measured the performance impact of the -fno-strict-aliasing flag? How much real-world performance are we really gaining from all this mess?

sestep1y ago

sapiogram1y ago

It's paywalled, is there a way to actually read it?

gpderetta1y ago

Depends a lot on the application. For many it matters little, but for some (mostly numerical), it can matter a lot.

twoodfin1y ago

Obviously it’s going to vary from program to program. And you always have to be skeptical that removing the safety for performance hasn’t given you a faster but faulty program.

That being said, my intuition matches what little anecdotal data I’ve seen from real perf-sensitive systems, and I’d ballpark 10-15% where it matters.

lmm1y ago

Real-world performance: not enough to be measurable, certainly remotely enough to make up for the time we lose to debugging.

But no-one cares about real-world performance, people pick C and pick a C compiler because they want the thing that's fastest on artificial microbenchmarks.

73kl4453dz1y ago

Lack of aliasing was historically fortran's advantage over C.

RossBencina1y ago· 5 in thread

Interesting that the article doesn't even entertain the obvious solution: remove strict aliasing requirements from the standards.

throw161803391y ago

There are no obvious changes to a widely used programming language standard.

gpderetta1y ago

If we can't even get that, I doubt strict aliasing will ever be voted out.

bluGill1y ago

zokier1y ago

on the other hand, it does say

> If I were writing correctness-oriented C that relied on these casts I wouldn’t even consider building it without -fno-strict-aliasing.

Arch-TK1y ago

"correctness-oriented C" definitionally cannot consider "[relying] on those casts".

icedchai1y ago· 2 in thread

I learned C on the Amiga, back in the late 80's, and the OS made heavy use of "OO-ish" physical subtyping with structs everywhere. I don't think anybody even thought about strict aliasing violations.

ajross1y ago

We're still waiting, alas.

flohofwoe1y ago

dang1y ago

Discussed at the time:

The Strict Aliasing Situation Is Pretty Bad - https://news.ycombinator.com/item?id=11288665 - March 2016 (67 comments)

j / k navigate · click thread line to collapse