There’s an interesting comment thread (including potential downsides) about this technique in the submission for “How to speed up the Rust compiler in September 2026” from a couple days ago: https://news.ycombinator.com/item?id=49923594
I wonder how this works when a crate's type checking does indeed depend on the bodies of one of its dependency. For example `-> impl Trait` return types famously leak the `Send`/`Sync`-ness of the type, so for type checking their uses you first need to resolve the concrete type behind and hence type checking the body of the function, which seems to go against the core idea of this proposal.
I guess you could type check the bodies of only these functions, or duplicating this work in the dependencies, but this will likely have further complications especially since this will also apply to all `async fn`s.
- It would have been very annoying if it didn't work like this
- Auto-traits also leak through private fields, so you would have the same issue if you e.g. wrapped the concrete return type in a new type and returned that one instead.
Talking with them presently! This specific set of changes won't get in (LLM written, little to no thinking through of the broader design), but I am hoping something like it shows up sometime.
Never thought of it like this, that synthetic code doesn't need to get into a codebase for it to benefit from the discussion around what was encoded with LLM.
Yes it's a pretty well thought out policy. It feels bad for someone to put in a month of work and have their code ultimately rejected, but if an llm does something sloppy with some light guidance in a few days, then great, experiment with some weird stuff and see what happens, then once you know it probably should work, go in for that month of work and make it happen.
Sucks when four out of twelve are flawed and you rule out the top 2 best solutions and 5 useful test cases you could learn from by just spray and praying... But at least it was fast!?
Yeah, that's not the way... It's synthesizing a single hypothesis at a time with attention to the emerging architecture, then decide to reimolement manually or reapply agents to refactor the code until it's reasonable.
That's a good point. I've gotten a few PRs to a project I work on, where the general idea is good, bit the actual code has serious issues, so I end up just implementing the idea myself in a better way.
It's a decent idea until 12000 bots start spamming a repo with janky code and poorly founded really complex suggestions for cs students to spray and pray resume padding.
As much as I am grateful for the stability that the rust stabilization process usually brings, I too am looking forward to seeing this in the mainline compiler in 2045.
Huh, I kinda thought something like this was already being done, but perhaps it was a bit later stage. Neat!
A kind of related shower thought I had, could rustc not instantiate generic methods, but write out what it expects to find (std::Vec<String>::push), and a separate build process listens to these, but makes sure already built ones are not built multiple times from the whole build.
I don’t actually know how much of duplication there usually is, but I had gotten the impression that it would be part of the problem.
Of course I’m no compiler engineer, and I’m sure there are complications like per crate build profiles, etc.
> I kinda thought something like this was already being done
It is. When building a crate, Cargo doesn't need to wait for all of that crate's dependencies to be completely finished compiling, instead it only needs to wait for each dependency to produce its metadata. The prototype here builds upon that existing machinery by causing metadata to be written earlier than it would otherwise be, prior to complete typechecking. It has the potential to increase the level of effective parallelism, so whether or not this has a benefit for a particular crate graph will depend on whether or not all of your cores are already effectively occupied during compilation or whether Cargo is letting cores lie idle while waiting on metadata to be produced.
Yes, if you mostly have no extra cores already or if you have very wide unrelated deps, this mostly does nothing. Typically, this is not the case though so there is some good perf gains.
From what I understood it's more like trust the external API of functions and continue with stuff that is downstream from that in parallel. If the function body then fails to type check then the compilation can be aborted and then you've done extra work, but that's a pretty rare case afaik.
(Copying and pasting from what I posted in rust zulip)
There is an .early-rmeta file that is generated that contains the result of type checking the outer part of the functions (arguments, returns) that is then passed down to all the other dependencies that can use it and get a "headstart" on their work. -Zearly-metadata is the flag for Rustc here to turn that on. It also needed to learn how to swap the real metadata in for the early stuff that it was previously using, once the real stuff showed up.
Rustc needs to be able to pause and resume at certain points and Cargo needs to be know how to look for and react to that, which is what -Zheadstart does for Cargo. So there's some changes job_queue that happen to there. I haven't checked this directly yet, but my understanding is that it doesn't kill the process, it just has it sit and wait, which normally would occupy a slot, but doesn't any longer.
---
So not a cache, just being able to pause and start work and better utilize the available slots.
Sounds like if this was incorporated you could include metadata files in crates and even avoid that analysis cost being repeated? Something like typescripts.d.ts files?
It’s a neat idea. My impression is that it would go against the grain of the compiler currently. The rmeta file and lots of the other rust outputs aren’t meant to be that stable that it could be portable like that. I’ll certainly mess around with this some and see if anything comes of it though!
It's deeper pipelining of the build units (similar to how CPU does pipelining). Rust has had pipelined builds for a while, but this extends it further.
As an aside, it's just silly to make the repository a bunch of checked-in patches. Git already does revision control. You don't need to do revision control in your revision control.
I wanted everything in one repo and this worked well enough. For presenting to rust compiler team I forked the repos in question and made branches for review with the patches.
Seems like an easy win! I'm kind of surprised nobody did this already. I guess someone will need to reimplement this by hand given Rust's AI policy, but this is still great because it demonstrates that it's a really good idea.
Rust has been taking steps in this direction for several years now, having first added pipelining via eager metadata emission in 2019: https://internals.rust-lang.org/t/evaluating-pipelined-rustc... . There have been several other steps in this process since then, e.g. working to make metadata less verbose, experimenting with different approaches to compression, etc. People have been eyeing making pipelining even more eager for a while now, as the submitter here noted a few days ago: https://news.ycombinator.com/item?id=49924257 , but there remain some decisions to be made regarding what to do when encountering errors in crates that were speculatively approved.
Rust's AI policy does not strictly forbid AI generated code, but it does require it to be used in a kind of responsible way. What that means though it pretty abstract even to me.
As for whether this change will make it in, the devil is as always in the details and will depend on whether this solution can work for all the edge cases in the language.
There’s an interesting comment thread (including potential downsides) about this technique in the submission for “How to speed up the Rust compiler in September 2026” from a couple days ago: https://news.ycombinator.com/item?id=49923594
The commenter there is the OP, for what it's worth
Additional information about potential downsides and prior art (including links to the previous time that this was attempted, in a slightly different form): https://github.com/PowderworksCode/headstart/blob/main/docs/...
I wonder how this works when a crate's type checking does indeed depend on the bodies of one of its dependency. For example `-> impl Trait` return types famously leak the `Send`/`Sync`-ness of the type, so for type checking their uses you first need to resolve the concrete type behind and hence type checking the body of the function, which seems to go against the core idea of this proposal.
I guess you could type check the bodies of only these functions, or duplicating this work in the dependencies, but this will likely have further complications especially since this will also apply to all `async fn`s.
Good question! They get marked as such and then typechecked as part of the early work before metadata is written/sent out. If they error out, crate fails, nothing happens, same as today. In practice, on the crates I was testing, there was more "regular" functions than these, so we could still get some good perf gains. https://github.com/PowderworksCode/headstart/blob/1c9d5cd691... https://github.com/PowderworksCode/headstart/blob/1c9d5cd691...
> For example `-> impl Trait` return types famously leak the `Send`/`Sync`-ness of the type,
Doesn't that mean that the return type isn't actually fully specified? That sounds wild to me. How did they end up there?
I'm not defending it (or claiming to really understand anything here), but this is where that came from:
https://rust-lang.github.io/rfcs/1522-conservative-impl-trai...
(OIBIT = auto trait, like Send/Sync)
Two reasons:
- It would have been very annoying if it didn't work like this
- Auto-traits also leak through private fields, so you would have the same issue if you e.g. wrapped the concrete return type in a new type and returned that one instead.
Excited to see if there is a path to getting this in the mainline compiler
Talking with them presently! This specific set of changes won't get in (LLM written, little to no thinking through of the broader design), but I am hoping something like it shows up sometime.
Never thought of it like this, that synthetic code doesn't need to get into a codebase for it to benefit from the discussion around what was encoded with LLM.
Yes it's a pretty well thought out policy. It feels bad for someone to put in a month of work and have their code ultimately rejected, but if an llm does something sloppy with some light guidance in a few days, then great, experiment with some weird stuff and see what happens, then once you know it probably should work, go in for that month of work and make it happen.
Using LLMs to build dozens of prototypes and test them before committing to a direction is one of their great strengths.
You no longer have to guess which paths will pay off. Have the LLM test them all, then go back and write the best one how you wanted to.
Sucks when four out of twelve are flawed and you rule out the top 2 best solutions and 5 useful test cases you could learn from by just spray and praying... But at least it was fast!?
Yeah, that's not the way... It's synthesizing a single hypothesis at a time with attention to the emerging architecture, then decide to reimolement manually or reapply agents to refactor the code until it's reasonable.
This is why LLM generated PRs are great… they should just be opened as issues instead.
That's a good point. I've gotten a few PRs to a project I work on, where the general idea is good, bit the actual code has serious issues, so I end up just implementing the idea myself in a better way.
It's a decent idea until 12000 bots start spamming a repo with janky code and poorly founded really complex suggestions for cs students to spray and pray resume padding.
A bit of KYC goes a long the way
As much as I am grateful for the stability that the rust stabilization process usually brings, I too am looking forward to seeing this in the mainline compiler in 2045.
Huh, I kinda thought something like this was already being done, but perhaps it was a bit later stage. Neat!
A kind of related shower thought I had, could rustc not instantiate generic methods, but write out what it expects to find (std::Vec<String>::push), and a separate build process listens to these, but makes sure already built ones are not built multiple times from the whole build.
I don’t actually know how much of duplication there usually is, but I had gotten the impression that it would be part of the problem.
Of course I’m no compiler engineer, and I’m sure there are complications like per crate build profiles, etc.
> I kinda thought something like this was already being done
It is. When building a crate, Cargo doesn't need to wait for all of that crate's dependencies to be completely finished compiling, instead it only needs to wait for each dependency to produce its metadata. The prototype here builds upon that existing machinery by causing metadata to be written earlier than it would otherwise be, prior to complete typechecking. It has the potential to increase the level of effective parallelism, so whether or not this has a benefit for a particular crate graph will depend on whether or not all of your cores are already effectively occupied during compilation or whether Cargo is letting cores lie idle while waiting on metadata to be produced.
Yes, if you mostly have no extra cores already or if you have very wide unrelated deps, this mostly does nothing. Typically, this is not the case though so there is some good perf gains.
So this is like a cache? Reminded me of https://turborepo.dev for TS, is it a similar concept?
From what I understood it's more like trust the external API of functions and continue with stuff that is downstream from that in parallel. If the function body then fails to type check then the compilation can be aborted and then you've done extra work, but that's a pretty rare case afaik.
ah ok, so more like an async type checking
It almost sounds like optimistic branch prediction when you put it that way
(Copying and pasting from what I posted in rust zulip)
There is an .early-rmeta file that is generated that contains the result of type checking the outer part of the functions (arguments, returns) that is then passed down to all the other dependencies that can use it and get a "headstart" on their work. -Zearly-metadata is the flag for Rustc here to turn that on. It also needed to learn how to swap the real metadata in for the early stuff that it was previously using, once the real stuff showed up.
Rustc needs to be able to pause and resume at certain points and Cargo needs to be know how to look for and react to that, which is what -Zheadstart does for Cargo. So there's some changes job_queue that happen to there. I haven't checked this directly yet, but my understanding is that it doesn't kill the process, it just has it sit and wait, which normally would occupy a slot, but doesn't any longer. ---
So not a cache, just being able to pause and start work and better utilize the available slots.
Sounds like if this was incorporated you could include metadata files in crates and even avoid that analysis cost being repeated? Something like typescripts.d.ts files?
It’s a neat idea. My impression is that it would go against the grain of the compiler currently. The rmeta file and lots of the other rust outputs aren’t meant to be that stable that it could be portable like that. I’ll certainly mess around with this some and see if anything comes of it though!
It's deeper pipelining of the build units (similar to how CPU does pipelining). Rust has had pipelined builds for a while, but this extends it further.
Woah nice thought! :) good idea
Those are big improvements!
As an aside, it's just silly to make the repository a bunch of checked-in patches. Git already does revision control. You don't need to do revision control in your revision control.
I wanted everything in one repo and this worked well enough. For presenting to rust compiler team I forked the repos in question and made branches for review with the patches.
Disagree. This was incredibly easy to test in a Bazel monorepo due to the ability to inject patch files into 3rd party dependencies (https://bazel.build/rules/lib/globals/module#parameters-13).
I would be very curious to know how that test went please!
Seems like an easy win! I'm kind of surprised nobody did this already. I guess someone will need to reimplement this by hand given Rust's AI policy, but this is still great because it demonstrates that it's a really good idea.
Rust has been taking steps in this direction for several years now, having first added pipelining via eager metadata emission in 2019: https://internals.rust-lang.org/t/evaluating-pipelined-rustc... . There have been several other steps in this process since then, e.g. working to make metadata less verbose, experimenting with different approaches to compression, etc. People have been eyeing making pipelining even more eager for a while now, as the submitter here noted a few days ago: https://news.ycombinator.com/item?id=49924257 , but there remain some decisions to be made regarding what to do when encountering errors in crates that were speculatively approved.
Uhm... emit the errors and fail the compilation? What else would you do?
Maybe there's some work to do to make that happen cleanly but the desired behaviour seems pretty obvious.
Rust's AI policy does not strictly forbid AI generated code, but it does require it to be used in a kind of responsible way. What that means though it pretty abstract even to me.
As for whether this change will make it in, the devil is as always in the details and will depend on whether this solution can work for all the edge cases in the language.