Rendered at 05:48:25 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
pizlonator 34 minutes ago [-]
It’s cool that this mentions Fil-C but it also undersells it. TFA also undersells CHERI. Fil-C doesn’t just “find a lot of temporal-safety” bugs. It closes off memory safety bugs (special and temporal) for exploit writers and ascribes a tight semantics to the whole language. CHERI makes some different trade offs but also gives a tight enough semantics that memory safety exploits aren’t going to work. Both CHERI and Fil-C are more comprehensive than Rust, since they attack the problem at the ABI level (and so you don’t get the problem that the protection only applies to the parts that were rewritten in the safe subset of a new language). Rust could be claimed to be better in that its compile time, but that doesn’t make a significant difference if you’re worried about the definedness of semantics or exploitability.
afdbcreid 17 minutes ago [-]
Both only attach provenance to allocations. The common example is:
struct User {
char name[100];
bool is_admin;
};
Where a buffer overflow can still overwrite `is_admin`.
Both also require recompilation of everything, which might be possible for CHERI but not for Fil-C - which is why, for example, there cannot be Fil-C support for Windows or macOS.
hn_submit 1 hours ago [-]
I believe this is barking up the wrong tree since IMHO C is just "high level assembly" for systems programming. As soon as you add runtime behavior to combat Undefined Behavior (UB) you're blowing up execution times. And static analysis can only go so far without blowing up compile times.
C is "the right tool for the right job" which is operating systems and its code which is called thousands of times per second. You cannot afford even one iota of runtime checks in that code. The developer must know what he's doing or he should get out of the kitchen.
We should discourage the usage of C in application programming and prod developers towards memory safe languages like Rust or Go.
And I'm not even sure if Rust solves this case as far as UB is concerned.
mdspan 4 minutes ago [-]
Operating systems code can definitely afford runtime checking if it's not too expensive in terms of performance. The Linux kernel has a WARN_ON macro that does exactly this. The tradeoff is worth it in a lot of cases if you're exchanging a small amount of performance for greater debug-ability or security.
gizmo686 42 minutes ago [-]
A high level assembly would not have a UB problem, because the generated machine code would closely map to your source code, so even technically undefined behaviour would end up doing the expected thing for the given hardware.
The problem with C is that modern compilers do a lot of transformations between your source code and the final machine code, so the actual behavior could be very far afield from what you would expect.
> And I'm not even sure if Rust solves this case as far as UB is concerned.
If your entire program is inside unsafe, then Rust is actually worse than C as far as UB is concerned. On the other hand, no one writes Rust like that, and Rust restricts all UB to unsafe blocks.
SkiFire13 11 minutes ago [-]
> C is just "high level assembly" for systems programming
It is not, and that's the issue. If it was just "high level assembly" there would be no UB, and no need to have UB. Instead you have UB (although arguably not all of it is really needed) because you need optimizations, which in turn you need because otherwise C would be too slow for that "operating system" job.
> code which is called thousands of times per second
Scripting languages can easily have loops running thousand of times per second and even more. You're off by some order of magnitudes here if you want to describe operations that happen in operating systems.
chasil 46 minutes ago [-]
C is also famously bad at floating point optimization.
Fortran has historcally led this realm (see the Numerical Recipes book).
Julia is a newer option, and I understand that both are commonly used in Python objects.
"Read the older 2nd ed. book in Fortran online for free."
jabl 12 minutes ago [-]
There are several reasons why Fortran historically has been faster. Many of these are because Fortran is less strict about what the compiler can do so allows optimizations that wouldn't be allowed in C. Such as:
- Procedure arguments are not allowed to alias. Similar to restrict pointers in C99+. This is often critical to allow loops to be vectorized, but the onus is on the programmer to ensure no aliasing or else you get UB.
- Unspecified evaluation order for expressions. E.g. C requires that "a+b+c+d" be evaluated as "((a+b)+c)+d)" and with floating point it can't do it another way due to rounding. Fortran can do e.g. "(a+b) + (c+d)" where each subterm can be computed in parallel, but again at the cost of slightly different results due to rounding behavior for floats.
- Old school Fortran lacked pointers which led programmers to program algorithms using arrays rather than fancier data structures, which cpu's love.
In principle there's nothing preventing a competent C or C++ programmer can reach Fortran level performance. In practice, might be difficult.
Of course, nowadays performance is much about designing for cache hierarchies (see e.g. "Data Oriented Design") where Fortran doesn't have a built-in advantage.
mianos 29 minutes ago [-]
There is not a lot of floating point in an operating system.
The issue with floating point and aliasing preventing vectorisation was from the late 1980s when early C compilers lacked sophisticated alias analysis and standards were loose. is not a really a thing anymore. C can go as fast, specially compiled with strict aliasing. Maybe more work in compilation.
chasil 3 hours ago [-]
"Some Honeywell machines, for example, had nine-bit bytes."
OS 2200 has 36-bit words. It is still a supported platform.
Here, we see that q[0] is an allocated but undefined byte. As per C99, this results in undefined behavior, however 20 years ago this was a good trick to get kinda-randomish bytes to use as a possible entropy source.
Someone claimed that the above brick of code will compile in newer versions of clang such that, since the complex cryptographic pseudo random number generator code depends on uninitialized but allocated memory, the entire cryptographic operation isn’t performed.
So I tested it against multiple versions of GCC and clang; I also tested it against TCC for good measure.
In all cases, with all levels of optimization, the cryptographic routine ran. I even ran it against clang 23. In cygwin, it was a randomish but consistent byte (except for clang at a higher level of optimization, at which point the uninitialized byte had a value of 0); in Ubuntu 26, the uninitialized memory consistently had a value of 0 (in tcc/gcc/clang).
I am hoping the up and coming C2y spec has very clearly defined behavior when using unintialized memory (ideally where it will work but the bytes can have any values).
Naturally, I have updated my code to no longer use uninitialized memory as a source of entropy. 20 years ago, MacOS didn’t support clock_gettime() with nanosecond resolution, so that wasn’t a portable way to get pseudo-random bits; these days clock_gettime() is universal across modern development environments, and it provides pretty good entropy (along with using /dev/urandom in *NIX, which isn’t in POSIX but is widely supported, as well as CryptGenRandom() in the legacy Win32 port). [2]
[1] Said person said the appendices to C99 aren’t authoritative, but if something is in the spec, including in the appendices, it’s authoritative.
[2] I don’t blindly trust /dev/urandom to always make really hard to guess pseudo-random bits, because my code is open source, and, as such, doesn’t just compile in Linux. It often times will be compiled in embedded systems, and even Linux has had at times issues with /dev/urandom on Raspberry Pis.
[3] I would also like to see uint8_t, int8_t, uint16_t, int16_t, uint32_t, int32_t, uint64_t, and int64_t mandated. They exist in C99, but aren’t mandated, even though every real world compiler from this century supports all of the above types. Yes, I know about _BitInt(8/16/32/64/128/etc.) but a compiler from 2004—and yes I still use one to make win32 binaries—doesn’t support these new C23 datatypes.
jcranmer 23 minutes ago [-]
> [1] Said person said the appendices to C99 aren’t authoritative, but if something is in the spec, including in the appendices, it’s authoritative.
That's not true. There is a difference between normative text and informative text. Informative text is not authoritative, and you were citing an appendix that is labeled as informative.
> I am hoping the up and coming C2y spec has very clearly defined behavior when using unintialized memory (ideally where it will work but the bytes can have any values).
It won't. Uninitialized memory can't have "very clearly defined behavior" without breaking essentially every single implementation, and WG14 is very loth to break existing implementations.
afdbcreid 8 minutes ago [-]
> It won't. Uninitialized memory can't have "very clearly defined behavior" without breaking essentially every single implementation, and WG14 is very loth to break existing implementations.
You can initialize with some bit pattern. This is what was accepted for C++ (but only for certain types of memory).
int* p = malloc(sizeof(int));
int v = *p;
if (!(v < 0 || v == 0 || v > 0)) {
exit(1);
}
This compiles into `exit(1)` in Clang under -O3.
fragmede 15 minutes ago [-]
Interesting. x86_64 has RDRAND for entropy (and RDSEED to seed). Oh I guess there's also ARM these days but I bet they got one too.
strenholme 5 minutes ago [-]
I write code which runs in embedded spaces, so a lot of ARM and even RISC-V. So I have a simple, portable source of pseudo-random numbers which will give good random numbers on a potato, which means rolling my own.
fithisux 17 minutes ago [-]
We need to learn to use assembly and call it from C. C was never meant to be the end of programming. Reducing undefined behavior is the least the standards body can do. Fil-C is also an extremely useful tool.
The real problem I see is the inability of independent compiler writers upgrading to the latest standard. I think this is also an issue that the standards body should pay attention. Help implementers.
p1necone 3 hours ago [-]
The concept of undefined behaviour specific to C/C++ has always seemed batshit insane to me, and I'm yet to read anything about it that has made it seem any less so.
nananana9 3 hours ago [-]
It doesn't make sense to define what happens when you e.g. read from NULL because it's hardware specific - if you have virtual memory of some sort, you'd probably get a page mapping error. If you don't (embedded, WASM), you'd read back whatever value is at that address.
Does Rust define what I get when I dereference NULL in unsafe code? I doubt it, since it would require a NULL check before every pointer dereference.
The only insane thing about UB is that compiler writers took what everyone understand meant "the compiler emits what it emits and you get what you get" and turned it into "since it's undefined it means it can never happen so we can delete your null check".
eru 2 hours ago [-]
> It doesn't make sense to define what happens when you e.g. read from NULL because it's hardware specific - if you have virtual memory of some sort, you'd probably get a page mapping error. If you don't (embedded, WASM), you'd read back whatever value is at that address.
That's an argument in favour of 'implementation defined behaviour'. Not 'undefined behaviour'.
tczMUFlmoNk 2 hours ago [-]
I don't disagree with your general thesis, but I don't think it's right to say that defining behaviors like dereferencing a null would require a null check before every dereference. For example, Java defines the behavior of dereferencing null—it throws a `NullPointerException`. My understanding is that JVMs implement this by (a) representing Java `null` references as the zero pointer, (b) mapping the zero page with write permission disabled, so that accesses are guaranteed to segfault, and (c) trapping `SIGSEGV` to translate that back into a Java exception.
So, while it's accurate to say that inserting null checks before every dereference is one way that you could implement this to make it well-defined, that is not the only way. We have lots of clever tricks to solve problems more efficiently than may seem possible at first glance—Fil-C is a bit of a modern marvel in that regard!
nananana9 1 hours ago [-]
> My understanding is that JVMs implement this by (a) representing Java `null` references as the zero pointer, (b) mapping the zero page with write permission disabled
If you want to run everywhere where C does, you can't rely on that. I gave WASM as an example - that's a widely used target that just exposes a flat memory model where 0 is a valid address (unless they've released extensions I'm unaware of). Same deal with microcontrollers.
I can't think of a way you'd implement null trapping efficiently on those platforms.
charleslmunger 2 hours ago [-]
Sure. But consider what would happen if you had an array, and you looked up the nth element. If that base address is a null pointer and the index is greater than the size of your zero page reservation, it'll get some other address which is holding stuff. There's ways to deal with that too, of course, but not for free and it carries implications for other things. In Java this is avoided because arrays carry their length at the beginning, and you check that first for bounds, so if the array pointer was null you'd fault a small number of bytes past 0 and it still works.
Fil-C is amazing and a prime example that undefined behavior means implementor freedom, and the implementor can choose to always trap on null pointer use. Sometimes the implementor freedom doesn't buy you much; for example why should it be UB to do
(const char*)NULL + 1
Dereferencing null is and should be UB but why is just calculating a pointer problematic? I just did some research and some old architectures would actually trap on creating an invalid address. So if we want C to support those machines, the standard can't define the behavior to do something other than what the hardware does.
eru 1 hours ago [-]
That's another argument in favour of implementation defined behaviour, not undefined behaviour.
Btw, a pointer in C doesn't necessarily need to mean an address (invalid or not) in your underlying machine. C is a formally defined abstract language, not portable assembly.
afdbcreid 2 hours ago [-]
> The only insane thing about UB is that compiler writers took what everyone understand meant "the compiler emits what it emits and you get what you get" and turned it into "since it's undefined it means it can never happen so we can delete your null check".
IMO UB as "undefined but don't be crazy please" was the original meaning of the standard but people argue on that. It is a fact that compilers didn't exploit UB as strongly back then. However there is a good reason for this change: if you want formal semantics (which you do want, at least possibly) it is pretty much impossible to distinguish the two. If "undefined behavior" is undefined in the math sense, or in formal semantics of languages - the operation can reach any Abstract Machine state, then the fact that you cannot reason about anything follows immediately. The only dubious thing is time-travel, and this was indeed removed in the last version of the standard (and also for Rust now).
Veserv 2 hours ago [-]
Undefined behavior is just what happens when you violate a mandatory precondition.
if (x > 0) {…}. But what if you entered the body when x <= 0?
if (false) {…}. But what if you execute the body?
These are “impossible”. What happens when the impossible occurs is “undefined”.
When the older standards said signed integer overflow for addition is undefined what they are actually saying is that the real definition of + is:
int +(int x, int y) {
assert(in_range(actual_math_add(x, y), signed_int_min, signed_int_max));
return machine_add(x, y);
}
So of course what happens when you get signed overflow is undefined; you should hit that assert and your program should explode and die. You should “never” get to the next instruction.
But, in the interest of performance, “release mode” (which in this case is just any compilation) elides asserts since as a programmer you should not write code that asserts in much the same way that you should not write assert(false) in a normal code path that is supposed to run. Assertions are intended for “impossible” code paths and usually get compiled out in “release mode” though maybe your code is buggy and can actually hit them and then your program goes off the rails because it had a bug.
Put another way, if you did write assert(false) in a regular code path, would you find it unreasonable for the compiler to just delete the code after it? That is what undefined behavior is for.
drdexebtjl 1 hours ago [-]
This isn’t a good mental model. The compiler isn’t limited to just affecting the code after the undefined behavior. It is allowed to assume undefined behavior never happens for the entire program.
For example, this code with an improper guard:
if (!p) puts("error");
printf("%d", *p);
Since the program dereferences p in line 2, and dereferencing null is UB, the compiler is allowed to assume p is never null, so it’s allowed to delete line 1, even though it would have executed before the point where UB would happen.
Even worse, the compiler isn't just allowed to not do things you told it to do, it's also allowed to do anything too.
Veserv 39 minutes ago [-]
The question was: “The concept of undefined behaviour specific to C/C++ has always seemed batshit insane to me”.
I was explaining why undefined behavior as a concept is a very sensible idea. Whether the expansive interpretation of the optimizations you are allowed to do when encountering the “impossible” are reasonable is a different question.
drdexebtjl 24 minutes ago [-]
The part that is batshit insane is precisely that it makes the entire program impossible to reason about.
If UB meant what you described, it would be sensible, but it doesn't, and it's not.
mpyne 3 hours ago [-]
It's not specific to C or C++, though it is more prominent there.
Rust's unsafe mode, for instance, has undefined behavior. A whole list of them, in fact.
GrantMoyer 1 hours ago [-]
Let's take a simple example: writing past the end of an array. Allow me to argue with myself for a moment.
> Surely the compiler can just check if each access is valid.
Well, sometimes it can, but sometimes it doesn't know how long the array is. What if the array is passed as a pointer?
> Maybe each array could be annotated with its size at runtime, and accesses could be checked at runtime too.
That works, but it adds runtime cost that may legitimately be too much for some applications, for example, a Gameboy game (set aside that many Gameboy games were written in assembly).
> Fine, so we'll make the programmer promise to ensure array accesses are always valid. Maybe they'll make a mistake sometimes, but what's the worst that could happen? Throwing your hands in the air and saying the compiler is allowed to do anything, that's just stupid.
Well, maybe it's stupid, but this is one thing that could happen if you accidentally write past the end of an array: https://www.youtube.com/watch?v=Vjm8P8utT5g. I'm sure neither the programmers nor compiler writers intended that.
Ultimately, the compiler can't guarantee any behavior if its assumptions are violated. The example may seem contrived, but it demonstrates that, given the right circumstances, the results of the logical contradiction are unbounded. This is a direct consequence of the "Principle of explosion": https://en.wikipedia.org/wiki/Principle_of_explosion. On second thought, maybe the runtime costs of array bounds checking are an acceptable trade-off after all.
chasil 3 hours ago [-]
It was written for a PDP-11 with 64k of RAM.
There wasn't room for safe programming practices, and direct manipulation of the hardware was a design requirement.
It assumes that you know what you are doing.
There are also ports to the Zilog Z80, an architecture with similar limitations (UZI, FUZIX).
elais-dev 3 hours ago [-]
i've seen static analyzers catch many ub patterns, but guaranteeing zero ub needs whole‑program analysis that blows up compile time and still produces false positives that drown developers
eru 2 hours ago [-]
Yes, it's pretty much impossible to prove anything about arbitrary programs.
However if you are willing to restrict what programs you allow, you can make guarantees possible.
Silly example: if you compile valid (safe) Rust programs to C, you know that the resulting code will not invalidate Rust's borrowing rules by construction; and in principle you could try to establish this guarantee just from the C code alone, never having seen the Rust original.
However, you still wouldn't be able to have an algorithm that tells you for any arbitrary C code whether it has these problems or not.
robotresearcher 1 hours ago [-]
> if you compile valid (safe) Rust programs to C, you know that the resulting code will not invalidate Rust's borrowing rules
If you know the compiler is correct, which you don't.
42 minutes ago [-]
BeaverGoose 3 hours ago [-]
Make signed overflow defined please.
_kst_ 2 hours ago [-]
I disagree, though I wouldn't mind adding a mechanism to say that you want signed overflow to be well defined.
C23 already requires 2's-complement representation for signed integer types, but signed overflow still has undefined behavior. I think that mandating 2's-complement wraparound would be a mistake.
Some instances of undefined behavior can be detected at compile time. For example, if I write
int too_big = INT_MAX + 1;
a reasonably clever compiler can warn about it (and in fact both gcc and clang do so). If the result of INT_MAX + 1 were defined by the language to be INT_MIN, there would be no basis for such a warning.
If you evaluate n + 1 and it's possible for n to be equal to INT_MAX before the addition what do you want the result to be? Would quietly yielding INT_MIN really be useful?
Ideally, if I (accidentally) evaluate INT_MAX + 1, I'd like to be told that I've made a mistake. C doesn't have a good mechanism for doing so.
gcc has a non-standard option "-fsanitize=signed-integer-overflow" that can be used to catch signed overflow at runtime. If signed overflow yielded a well defined result, that option would be non-conforming.
eru 2 hours ago [-]
> a reasonably clever compiler can warn about it (and in fact both gcc and clang do so). If the result of INT_MAX + 1 were defined by the language to be INT_MIN, there would be no basis for such a warning.
Compilers warn about perfectly well defined behaviour all the time. That's why these are warnings, not errors.
afdbcreid 2 hours ago [-]
You can perfectly specify that integer overflow either traps or wraps-around (this is what Rust does). Then the flag would be conforming. You can also say it is Erroneous Behavior (a new term in C++, not yet adopted for C AFAIK) and specified to wrap-around, which will also allow the flag (and arguably models Rust behavior more closely).
creato 2 hours ago [-]
Why would this help? Unsigned integer overflow is defined behavior, that causes basically the same set of bugs in practice. In a way it is worse, because at least runtime UB checkers have a reason to complain about signed integer overflow, but they won't complain about unsigned integer overflow.
eru 2 hours ago [-]
It would be an improvement.
However if you wan, you can already get that via a flag in pretty much any C compiler you care about.
homosapien97 1 hours ago [-]
Use -fwrapv
pajko 51 minutes ago [-]
Or -ftrapv depending on usage, both converting UB to IB. Generally letting a number overflowing and wrapping around is opening the door to let in the bad things creeping in the shadow.
jdw64 3 hours ago [-]
[dead]
jdw64 3 hours ago [-]
If we follow this video, the real constraints essentially mean POSIX and ABI. This implies that the contract a programmer must understand is distributed across multiple layers—not just the language specification, but effectively the operating system and ABI as well.
Ultimately, this leads to the conclusion that even if the language itself is fundamentally free, constraints are inherently necessary at its lower layers.
If we were to bloat the compiler—that is, if we restricted freedom like Rust does with its borrow checker—then the freedom available to the programmer would vanish. If that happens, people might grow weary of a language that is supposed to be free.
In the end, any single layer inherently restricts freedom. In other words, I feel there is a need to transfer the complexity that a programmer must manage over to a mechanical management system, but which layer would be best for that?
Currently, based on experience, this level of complexity is categorized into the language layer, and that level of complexity into the operating system layer. But in the future, won't there be some sort of complexity theorem that determines which layer minimizes complexity the most, and won't systems be completely rewritten based on that?
Both also require recompilation of everything, which might be possible for CHERI but not for Fil-C - which is why, for example, there cannot be Fil-C support for Windows or macOS.
C is "the right tool for the right job" which is operating systems and its code which is called thousands of times per second. You cannot afford even one iota of runtime checks in that code. The developer must know what he's doing or he should get out of the kitchen.
We should discourage the usage of C in application programming and prod developers towards memory safe languages like Rust or Go.
And I'm not even sure if Rust solves this case as far as UB is concerned.
The problem with C is that modern compilers do a lot of transformations between your source code and the final machine code, so the actual behavior could be very far afield from what you would expect.
> And I'm not even sure if Rust solves this case as far as UB is concerned. If your entire program is inside unsafe, then Rust is actually worse than C as far as UB is concerned. On the other hand, no one writes Rust like that, and Rust restricts all UB to unsafe blocks.
It is not, and that's the issue. If it was just "high level assembly" there would be no UB, and no need to have UB. Instead you have UB (although arguably not all of it is really needed) because you need optimizations, which in turn you need because otherwise C would be too slow for that "operating system" job.
> code which is called thousands of times per second
Scripting languages can easily have loops running thousand of times per second and even more. You're off by some order of magnitudes here if you want to describe operations that happen in operating systems.
Fortran has historcally led this realm (see the Numerical Recipes book).
Julia is a newer option, and I understand that both are commonly used in Python objects.
https://numerical.recipes/
"Read the older 2nd ed. book in Fortran online for free."
- Procedure arguments are not allowed to alias. Similar to restrict pointers in C99+. This is often critical to allow loops to be vectorized, but the onus is on the programmer to ensure no aliasing or else you get UB.
- Unspecified evaluation order for expressions. E.g. C requires that "a+b+c+d" be evaluated as "((a+b)+c)+d)" and with floating point it can't do it another way due to rounding. Fortran can do e.g. "(a+b) + (c+d)" where each subterm can be computed in parallel, but again at the cost of slightly different results due to rounding behavior for floats.
- Old school Fortran lacked pointers which led programmers to program algorithms using arrays rather than fancier data structures, which cpu's love.
In principle there's nothing preventing a competent C or C++ programmer can reach Fortran level performance. In practice, might be difficult.
Of course, nowadays performance is much about designing for cache hierarchies (see e.g. "Data Oriented Design") where Fortran doesn't have a built-in advantage.
The issue with floating point and aliasing preventing vectorisation was from the late 1980s when early C compilers lacked sophisticated alias analysis and standards were loose. is not a really a thing anymore. C can go as fast, specially compiled with strict aliasing. Maybe more work in compilation.
OS 2200 has 36-bit words. It is still a supported platform.
https://en.wikipedia.org/wiki/UNIVAC_1100/2200_series
This platform was the first SMP UNIX implementation:
"Any configuration supplied by Sperry, including multiprocessor ones, can run the UNIX system."
https://www.nokia.com/bell-labs/about/dennis-m-ritchie/other...
Yes, but that's perhaps an argument for 'implementation defined behaviour', not in favour of 'undefined behaviour'.
1790673092 | Reducing undefined behavior in the C language | https://lwn.net/SubscriberLink/1095811/efcdbcf080cfa4c6/ | https://news.ycombinator.com/item?id=49890290 | 0 comments
As per the linked article:
“There are currently about 100 instances of undefined behavior in the C standard, but the in-progress C2y draft has removed 45 of them.”
I wonder how they handle the specific case of uninitialized but allocated memory.
Let’s look at something which will result in undefined behavior in C99: [1]
The key part of the above brick of code is this: Here, we see that q[0] is an allocated but undefined byte. As per C99, this results in undefined behavior, however 20 years ago this was a good trick to get kinda-randomish bytes to use as a possible entropy source.Someone claimed that the above brick of code will compile in newer versions of clang such that, since the complex cryptographic pseudo random number generator code depends on uninitialized but allocated memory, the entire cryptographic operation isn’t performed.
So I tested it against multiple versions of GCC and clang; I also tested it against TCC for good measure.
In all cases, with all levels of optimization, the cryptographic routine ran. I even ran it against clang 23. In cygwin, it was a randomish but consistent byte (except for clang at a higher level of optimization, at which point the uninitialized byte had a value of 0); in Ubuntu 26, the uninitialized memory consistently had a value of 0 (in tcc/gcc/clang).
I am hoping the up and coming C2y spec has very clearly defined behavior when using unintialized memory (ideally where it will work but the bytes can have any values).
Naturally, I have updated my code to no longer use uninitialized memory as a source of entropy. 20 years ago, MacOS didn’t support clock_gettime() with nanosecond resolution, so that wasn’t a portable way to get pseudo-random bits; these days clock_gettime() is universal across modern development environments, and it provides pretty good entropy (along with using /dev/urandom in *NIX, which isn’t in POSIX but is widely supported, as well as CryptGenRandom() in the legacy Win32 port). [2]
[1] Said person said the appendices to C99 aren’t authoritative, but if something is in the spec, including in the appendices, it’s authoritative.
[2] I don’t blindly trust /dev/urandom to always make really hard to guess pseudo-random bits, because my code is open source, and, as such, doesn’t just compile in Linux. It often times will be compiled in embedded systems, and even Linux has had at times issues with /dev/urandom on Raspberry Pis.
[3] I would also like to see uint8_t, int8_t, uint16_t, int16_t, uint32_t, int32_t, uint64_t, and int64_t mandated. They exist in C99, but aren’t mandated, even though every real world compiler from this century supports all of the above types. Yes, I know about _BitInt(8/16/32/64/128/etc.) but a compiler from 2004—and yes I still use one to make win32 binaries—doesn’t support these new C23 datatypes.
That's not true. There is a difference between normative text and informative text. Informative text is not authoritative, and you were citing an appendix that is labeled as informative.
> I am hoping the up and coming C2y spec has very clearly defined behavior when using unintialized memory (ideally where it will work but the bytes can have any values).
It won't. Uninitialized memory can't have "very clearly defined behavior" without breaking essentially every single implementation, and WG14 is very loth to break existing implementations.
You can initialize with some bit pattern. This is what was accepted for C++ (but only for certain types of memory).
The real problem I see is the inability of independent compiler writers upgrading to the latest standard. I think this is also an issue that the standards body should pay attention. Help implementers.
Does Rust define what I get when I dereference NULL in unsafe code? I doubt it, since it would require a NULL check before every pointer dereference.
The only insane thing about UB is that compiler writers took what everyone understand meant "the compiler emits what it emits and you get what you get" and turned it into "since it's undefined it means it can never happen so we can delete your null check".
That's an argument in favour of 'implementation defined behaviour'. Not 'undefined behaviour'.
(Here's a link with a bit more detail: https://courses.cs.vt.edu/cs3214/spring2026/questions/catchi...)
So, while it's accurate to say that inserting null checks before every dereference is one way that you could implement this to make it well-defined, that is not the only way. We have lots of clever tricks to solve problems more efficiently than may seem possible at first glance—Fil-C is a bit of a modern marvel in that regard!
If you want to run everywhere where C does, you can't rely on that. I gave WASM as an example - that's a widely used target that just exposes a flat memory model where 0 is a valid address (unless they've released extensions I'm unaware of). Same deal with microcontrollers.
I can't think of a way you'd implement null trapping efficiently on those platforms.
Fil-C is amazing and a prime example that undefined behavior means implementor freedom, and the implementor can choose to always trap on null pointer use. Sometimes the implementor freedom doesn't buy you much; for example why should it be UB to do
Dereferencing null is and should be UB but why is just calculating a pointer problematic? I just did some research and some old architectures would actually trap on creating an invalid address. So if we want C to support those machines, the standard can't define the behavior to do something other than what the hardware does.Btw, a pointer in C doesn't necessarily need to mean an address (invalid or not) in your underlying machine. C is a formally defined abstract language, not portable assembly.
IMO UB as "undefined but don't be crazy please" was the original meaning of the standard but people argue on that. It is a fact that compilers didn't exploit UB as strongly back then. However there is a good reason for this change: if you want formal semantics (which you do want, at least possibly) it is pretty much impossible to distinguish the two. If "undefined behavior" is undefined in the math sense, or in formal semantics of languages - the operation can reach any Abstract Machine state, then the fact that you cannot reason about anything follows immediately. The only dubious thing is time-travel, and this was indeed removed in the last version of the standard (and also for Rust now).
if (x > 0) {…}. But what if you entered the body when x <= 0?
if (false) {…}. But what if you execute the body?
These are “impossible”. What happens when the impossible occurs is “undefined”.
When the older standards said signed integer overflow for addition is undefined what they are actually saying is that the real definition of + is:
So of course what happens when you get signed overflow is undefined; you should hit that assert and your program should explode and die. You should “never” get to the next instruction.But, in the interest of performance, “release mode” (which in this case is just any compilation) elides asserts since as a programmer you should not write code that asserts in much the same way that you should not write assert(false) in a normal code path that is supposed to run. Assertions are intended for “impossible” code paths and usually get compiled out in “release mode” though maybe your code is buggy and can actually hit them and then your program goes off the rails because it had a bug.
Put another way, if you did write assert(false) in a regular code path, would you find it unreasonable for the compiler to just delete the code after it? That is what undefined behavior is for.
For example, this code with an improper guard:
Since the program dereferences p in line 2, and dereferencing null is UB, the compiler is allowed to assume p is never null, so it’s allowed to delete line 1, even though it would have executed before the point where UB would happen.Even worse, the compiler isn't just allowed to not do things you told it to do, it's also allowed to do anything too.
I was explaining why undefined behavior as a concept is a very sensible idea. Whether the expansive interpretation of the optimizations you are allowed to do when encountering the “impossible” are reasonable is a different question.
If UB meant what you described, it would be sensible, but it doesn't, and it's not.
Rust's unsafe mode, for instance, has undefined behavior. A whole list of them, in fact.
> Surely the compiler can just check if each access is valid.
Well, sometimes it can, but sometimes it doesn't know how long the array is. What if the array is passed as a pointer?
> Maybe each array could be annotated with its size at runtime, and accesses could be checked at runtime too.
That works, but it adds runtime cost that may legitimately be too much for some applications, for example, a Gameboy game (set aside that many Gameboy games were written in assembly).
> Fine, so we'll make the programmer promise to ensure array accesses are always valid. Maybe they'll make a mistake sometimes, but what's the worst that could happen? Throwing your hands in the air and saying the compiler is allowed to do anything, that's just stupid.
Well, maybe it's stupid, but this is one thing that could happen if you accidentally write past the end of an array: https://www.youtube.com/watch?v=Vjm8P8utT5g. I'm sure neither the programmers nor compiler writers intended that.
Ultimately, the compiler can't guarantee any behavior if its assumptions are violated. The example may seem contrived, but it demonstrates that, given the right circumstances, the results of the logical contradiction are unbounded. This is a direct consequence of the "Principle of explosion": https://en.wikipedia.org/wiki/Principle_of_explosion. On second thought, maybe the runtime costs of array bounds checking are an acceptable trade-off after all.
There wasn't room for safe programming practices, and direct manipulation of the hardware was a design requirement.
It assumes that you know what you are doing.
There are also ports to the Zilog Z80, an architecture with similar limitations (UZI, FUZIX).
However if you are willing to restrict what programs you allow, you can make guarantees possible.
Silly example: if you compile valid (safe) Rust programs to C, you know that the resulting code will not invalidate Rust's borrowing rules by construction; and in principle you could try to establish this guarantee just from the C code alone, never having seen the Rust original.
However, you still wouldn't be able to have an algorithm that tells you for any arbitrary C code whether it has these problems or not.
If you know the compiler is correct, which you don't.
C23 already requires 2's-complement representation for signed integer types, but signed overflow still has undefined behavior. I think that mandating 2's-complement wraparound would be a mistake.
Some instances of undefined behavior can be detected at compile time. For example, if I write
a reasonably clever compiler can warn about it (and in fact both gcc and clang do so). If the result of INT_MAX + 1 were defined by the language to be INT_MIN, there would be no basis for such a warning.If you evaluate n + 1 and it's possible for n to be equal to INT_MAX before the addition what do you want the result to be? Would quietly yielding INT_MIN really be useful?
Ideally, if I (accidentally) evaluate INT_MAX + 1, I'd like to be told that I've made a mistake. C doesn't have a good mechanism for doing so.
gcc has a non-standard option "-fsanitize=signed-integer-overflow" that can be used to catch signed overflow at runtime. If signed overflow yielded a well defined result, that option would be non-conforming.
Compilers warn about perfectly well defined behaviour all the time. That's why these are warnings, not errors.
However if you wan, you can already get that via a flag in pretty much any C compiler you care about.
Currently, based on experience, this level of complexity is categorized into the language layer, and that level of complexity into the operating system layer. But in the future, won't there be some sort of complexity theorem that determines which layer minimizes complexity the most, and won't systems be completely rewritten based on that?