Rendered at 20:51:50 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
nlehuen 8 hours ago [-]
> readable C
> return uu____0;
If readability is the selling point, why not go the extra millimeter and add a heuristic so that this variable is named return_value for instance?
taosher 4 hours ago [-]
[flagged]
arbor-group 7 hours ago [-]
[flagged]
weinzierl 8 hours ago [-]
I highly recommend checking out the other projects from AeneasVerif
(and Jonathan Protzenko).
Scylla is kind of Eurydice's dual. But also Charon and the eponymous Aeneas.
Also,
"There is ongoing work to integrate Eurydice-generated code for both Microsoft and Google’s respective crypto libraries." [1]
which I find incredibly cool. This link is also a brilliant intro to Eurydice.
If it's readable C, this sounds great for a full source bootstrap of the rust compiler.
knxsnsn 12 hours ago [-]
Why would it need to be readable to be applicable to bootstrapping?
stabbles 11 hours ago [-]
It depends entirely on your trust model. If you just want to be able to compile a rust compiler on a new platform that only has a C compiler, you wouldn't care. If you worry about a Thompson-style trojan horse, then you want to be able to audit all sources in the bootstrap chain.
Some idealistic projects like live-bootstrap go a step further and don't even allow generated sources in the full source bootstrap chain, even if they are readable in principle, because the xz debacle showed that a configure script, which is also readable in principle, can contain a backdoor. In those projects you could probably still bootstrap Eurydice from sources, and then do the rust -> c translation of the rust compiler as an ordinary build step, before compiling it with a bootstrapped C compiler.
david-gpu 10 hours ago [-]
> If you worry about a Thompson-style trojan horse, then you want to be able to audit all sources in the bootstrap chain.
Wouldn't you want to audit the final outputs? Except nobody is going to bother analyzing my hand all the assembly produced by the final tool.
In other words, I am not seeing the value of pursuing the discovery of Thompson-style Trojans, no matter what the original source language was. It is one of these risks that one has to live with.
vlovich123 5 hours ago [-]
It’s actually been solved how you can defeat it using “ diverse double-compiling”
* compile “real” compiler with a different trusted compiler, producing “independent real compiler”
* compile the “real” compiler again using itself and the “independent compiler” - if the outputs don’t match, then you have detected the issue.
Source level transformation is a brilliant way to generate the “independent compiler”.
You are qualified to have your opinions, but do you have the credentials?
peter_d_sherman 4 hours ago [-]
Stabbles gets it! (Also, thank you for your other excellent posts on HN -- you are clearly not a bot like 50% of the other accounts are!)
Now, trust/security is one aspect of things, that's true...
But another important aspect of all of this is building understandable, maintainable, controllable software and AI's from the ground up, for future generations...
Let's suppose that in the future, there is no longer source code for compilers, they are only black-box compilers that future AI (which is also black-box by then) uses... At that point, humanity has lost control! No more source can be compiled without a black-box compiler, without an AI, and any source code so compiled will be controlled by whatever unknown things have been implanted in the original compiler!
Phrased another way, when humanity requires AI to generate new AI's (i.e., no scientists, computer scientists, programmers or other people know how to generate an AI from the ground up, from transistors on up, and control every layer of abstraction on top of that from assemblers to compilers to neural networks)... then humanity is lost, because at that point, humanity (not knowing how to build AI again from scratch, and control it at every step) quite literally has lost control of AI!
Future humanity cannot have that.
Future humanity requires, absolutely requires, transparency at every abstraction level... no black boxes should exist anywhere!
Future humanity also requires the requisite education, to be able to understand, manage and control each level of hardware/software abstraction, but that is outside of the scope of this HN post... :-)
Anyway, Stabbles, you get it!
csmlab_notes 9 hours ago [-]
[dead]
Neywiny 20 hours ago [-]
I wish this covered more features specific to rust that make it more runtime safer not just compile time safer. I guess by the time it's IR it's the same but moving it back up to C would be nice to see. Like bounds checking.
Lvl999Noob 20 hours ago [-]
I think those are covered. Any checks that Rust adds will be present in the IR (either coded in the original source or implied through rust semantics). Those checks would then get put in the compiled C output. The ideal output of Eurydice would be to have exactly the same semantics and safety as provided by Rust. Since their use case is to use Rust to code on platform without support for a rust compiler, it is a rather important goal to have too.
Neywiny 9 hours ago [-]
Sorry not the software, I meant the article/write-up. I'm sure the software covers everything
fsloth 7 hours ago [-]
”This can result in several different implementations of a function that differ only by type — often, the more idiomatic C approach would be to use macros or void * arguments.”
”Idiomatic” - is a ridiculous tautology I wish people discussing languages would stop using.
The only non-idiomatic code written seriosly is that which the compiler/interpreter does not understand.
Claiming universal superiority of style is shallow and thoughtless.
Intentional obfuscation is a different thing.
genxy 1 hours ago [-]
s/idiomatic/normative
When computer people name things, it is with a tenuous grasp of meaning of the words they are using. You kinda have to roll with it.
fsloth 37 minutes ago [-]
Sure - when discussing C of all languages this expression felt shallower than usual though.
If you are really enthusiastic you can write _obsfuscated_ C. And people can kind of agree what that is when they see it.
But _idiomatic_ ? Good luck defining what that is.
The term has sort of stuck in Python and there it's purpose is a bit more clear. There is this _object_ way of doing things, there is this _expression_ way of doing things etc and the interaction patterns work really nicely.
But you can't really use that expression in other languages which lack the same level of institutional design without sounding very much out of touch.
chias 16 hours ago [-]
What an excellent choice of product name :D
"You're da C"
nfw2 15 hours ago [-]
"you read da c"
Only know the pronunciation due to hadestown
nfw2 15 hours ago [-]
Really more like "you're rid o(f) c" though, have you considered making it compile the other way
egeozcan 14 hours ago [-]
Sorry but even threads about fringe mathematics are somewhat easier to decipher for me than this thread.
ELI5, anyone?
Twisol 14 hours ago [-]
"Eurydice" is pronounced roughly like each of the ancestor posts. (It isn't "you-ree-dice".)
1718627440 8 hours ago [-]
Isn't "Eurydice" Greek, so I would expect it to be pronounced more like [eu̯.ry.dí.kɛː], so nothing remotely similar to "you"?
scns 10 hours ago [-]
> It isn't "you-ree-dice"
True, it is "you-ri-di-cee"
fnord77 14 hours ago [-]
the circle is complete
kevincox 9 hours ago [-]
I wonder if you keep compiling back and forth would you ever hit a cycle or would it be constantly growing?
freecodeio 2 hours ago [-]
theoretically it should be a circle, if it keeps constantly growing then you're looking at bugs
fithisux 15 hours ago [-]
D and C++ would benefit from something like this.
pjmlp 14 hours ago [-]
Why?
C++ started as a preprocessor that translated into C, nowadays all relevant C compilers are written in C++.
D has enough backeds already available.
uecker 13 hours ago [-]
People still maintain old versions of GCC just to maintain a bootstrap path.
I also think it would be helpful to how have a useful intermediate result to look at. So I agree, having C++ to C compiler would be useful.
pjmlp 12 hours ago [-]
Intermediate result is the compilation to Assembly source code, and C++ insights like tooling.
Those people should move on, there isn't a single relevant platform without at least C++98 compiler, and there are other ways to bootstrap compilers.
tialaramex 10 hours ago [-]
I don't see what's wrong with e.g. LLVM IR for that purpose.
C seems to me a really poor choice because it's barely more expressive than one of the machine code architectures and yet unlike them it's not even a real machine we can know about - some of the edge cases are just a shrug emoji.
The LLVM IR is buggy of course, and under-specified, but we could imagine improving those things more easily than fixing C for this purpose.
pjmlp 9 hours ago [-]
In fact, looking at several languages since Fortran, some kind of minimal bytecode was a common approach.
That is also how Pascal P-Code started, UCSD reused and extend it for UCSD Pascal, however Niklaus Wirth designed it originally to port Pascal compilers.
Then he did it again with Modula-2 (M-Code), and Oberon (slim binaries, although here the idea came from Michael Franz).
Naturally there were many others, and best of all no UB surprises with backend optimisations.
uecker 9 hours ago [-]
Nothing against byte code for implementation, but it is also not readable.
It is easy to avoid UB surprises in C when generating code, so I do not think there is any good argument against C at this point.
uecker 10 hours ago [-]
C is a lot more readable than assembly or LLVM IR, so no - do not agree it that is a poor choice. I also do not see the edge cases in practice.
tialaramex 8 hours ago [-]
Both you and I think C is more readable than LLVM IR but I'm not at all convinced that's just universally true -- I think it's because we're both experienced C programmers and so it's just our bias.
In particular I think LLVM's poison and undef values are valuable for understanding what some C++ means and why. Seeing that some C++ ends up as a select to choose between certain values, some of which are poison, makes it apparent why an optimisation choice is sound or not, where the likely C would leave this unsaid.
uecker 7 hours ago [-]
LLVM IR is very low-level and verbose and I think that makes it generally less readable. It is much like assembler. I have not seen people using it to write useful programs, and I think this would be painful.
I agree that undef / poison etc. are perhaps useful for optimization, but I was not proposing C as an intermediate language for optimization but as the result of C++ lowering. There are still some weird C++ semantics and specific features that would need extensions.
fithisux 11 hours ago [-]
For platforms not having a D compiler.
The C++ subset target is of theoretical interest and because some compilers cannot support a C++ compiler of the latest standard. OpenWatcom comes to mind.
Of course the question is which C++ standard in and what C standard out!
pjmlp 10 hours ago [-]
So target C++98, or C++ARM.
As for not having a compiler, do it like in the old days, cross compilation.
See Go and Zig for good examples of old practices brought into modern times.
IshKebab 9 hours ago [-]
Apprently https://edgcpp.org/ can do that for C++. I dunno why you'd need to though?
If readability is the selling point, why not go the extra millimeter and add a heuristic so that this variable is named return_value for instance?
Scylla is kind of Eurydice's dual. But also Charon and the eponymous Aeneas.
Also, "There is ongoing work to integrate Eurydice-generated code for both Microsoft and Google’s respective crypto libraries." [1] which I find incredibly cool. This link is also a brilliant intro to Eurydice.
[1] https://jonathan.protzenko.fr/2025/10/28/eurydice.html
Some idealistic projects like live-bootstrap go a step further and don't even allow generated sources in the full source bootstrap chain, even if they are readable in principle, because the xz debacle showed that a configure script, which is also readable in principle, can contain a backdoor. In those projects you could probably still bootstrap Eurydice from sources, and then do the rust -> c translation of the rust compiler as an ordinary build step, before compiling it with a bootstrapped C compiler.
Wouldn't you want to audit the final outputs? Except nobody is going to bother analyzing my hand all the assembly produced by the final tool.
In other words, I am not seeing the value of pursuing the discovery of Thompson-style Trojans, no matter what the original source language was. It is one of these risks that one has to live with.
* compile “real” compiler with a different trusted compiler, producing “independent real compiler”
* compile the “real” compiler again using itself and the “independent compiler” - if the outputs don’t match, then you have detected the issue.
Source level transformation is a brilliant way to generate the “independent compiler”.
https://en.wikipedia.org/wiki/Backdoor_(computing)#Counterme...
Now, trust/security is one aspect of things, that's true...
But another important aspect of all of this is building understandable, maintainable, controllable software and AI's from the ground up, for future generations...
Let's suppose that in the future, there is no longer source code for compilers, they are only black-box compilers that future AI (which is also black-box by then) uses... At that point, humanity has lost control! No more source can be compiled without a black-box compiler, without an AI, and any source code so compiled will be controlled by whatever unknown things have been implanted in the original compiler!
Phrased another way, when humanity requires AI to generate new AI's (i.e., no scientists, computer scientists, programmers or other people know how to generate an AI from the ground up, from transistors on up, and control every layer of abstraction on top of that from assemblers to compilers to neural networks)... then humanity is lost, because at that point, humanity (not knowing how to build AI again from scratch, and control it at every step) quite literally has lost control of AI!
Future humanity cannot have that.
Future humanity requires, absolutely requires, transparency at every abstraction level... no black boxes should exist anywhere!
Future humanity also requires the requisite education, to be able to understand, manage and control each level of hardware/software abstraction, but that is outside of the scope of this HN post... :-)
Anyway, Stabbles, you get it!
”Idiomatic” - is a ridiculous tautology I wish people discussing languages would stop using.
The only non-idiomatic code written seriosly is that which the compiler/interpreter does not understand.
Claiming universal superiority of style is shallow and thoughtless.
Intentional obfuscation is a different thing.
When computer people name things, it is with a tenuous grasp of meaning of the words they are using. You kinda have to roll with it.
If you are really enthusiastic you can write _obsfuscated_ C. And people can kind of agree what that is when they see it.
But _idiomatic_ ? Good luck defining what that is.
The term has sort of stuck in Python and there it's purpose is a bit more clear. There is this _object_ way of doing things, there is this _expression_ way of doing things etc and the interaction patterns work really nicely.
But you can't really use that expression in other languages which lack the same level of institutional design without sounding very much out of touch.
"You're da C"
Only know the pronunciation due to hadestown
ELI5, anyone?
True, it is "you-ri-di-cee"
C++ started as a preprocessor that translated into C, nowadays all relevant C compilers are written in C++.
D has enough backeds already available.
I also think it would be helpful to how have a useful intermediate result to look at. So I agree, having C++ to C compiler would be useful.
Those people should move on, there isn't a single relevant platform without at least C++98 compiler, and there are other ways to bootstrap compilers.
C seems to me a really poor choice because it's barely more expressive than one of the machine code architectures and yet unlike them it's not even a real machine we can know about - some of the edge cases are just a shrug emoji.
The LLVM IR is buggy of course, and under-specified, but we could imagine improving those things more easily than fixing C for this purpose.
That is also how Pascal P-Code started, UCSD reused and extend it for UCSD Pascal, however Niklaus Wirth designed it originally to port Pascal compilers.
Then he did it again with Modula-2 (M-Code), and Oberon (slim binaries, although here the idea came from Michael Franz).
Naturally there were many others, and best of all no UB surprises with backend optimisations.
It is easy to avoid UB surprises in C when generating code, so I do not think there is any good argument against C at this point.
In particular I think LLVM's poison and undef values are valuable for understanding what some C++ means and why. Seeing that some C++ ends up as a select to choose between certain values, some of which are poison, makes it apparent why an optimisation choice is sound or not, where the likely C would leave this unsaid.
I agree that undef / poison etc. are perhaps useful for optimization, but I was not proposing C as an intermediate language for optimization but as the result of C++ lowering. There are still some weird C++ semantics and specific features that would need extensions.
The C++ subset target is of theoretical interest and because some compilers cannot support a C++ compiler of the latest standard. OpenWatcom comes to mind.
Of course the question is which C++ standard in and what C standard out!
As for not having a compiler, do it like in the old days, cross compilation.
See Go and Zig for good examples of old practices brought into modern times.