When notebookLM was new, it was interesting to listen to the podcasts. Then the novelty wore off, and I wanted something where I can interact with the podcasters but it was janky as hell.
My current “audio-learning” hack is ChatGPT Live which has become shockingly good after being awful compared to Claude Voice (Let’s not even talk about Gemini voice which is still bad).
I go on a walk and dump a paper or article link in the chat, and ask chatGPT Live to walk me through the content in small nuggets, so I can discuss them interactively. For deeper topics I have it quiz me Socratic style so I’m not just passively listening, and actually thinking through problems or ideas.
There is actually an interrupt mode in NotebookLM now.
Overall I've found it the best AI product Google have. Only complaint I have about it is the hyper positive US corporate accents get pretty annoying pretty quickly.
The realtime voice in ChatGPT is excellent, the newer model is a big step up too.
I wonder how much of the annoying aspects of the podcasters can be controlled by the custom user instructions. I hate sitting through the patter of how "we're not political, just reporting the views of the sources."
"We have some REALLY interesting sources that we're discussing today"
Yeah I think the podcasts were just a clever hack until AI is fast enough to support realtime. It’s practically there now, so a fake podcast is a funny relic and hard to sit through.
Too much friction in notebookLM - several minutes to make a podcast, and they disappear after a while, and the interrupt feature is super-janky.
> Too much friction in notebookLM
This was my main problem with it. Loved it otherwise, but the lack of an API or other automation so I could dump stuff in it and have a podcast ready to download and listen to on my subway ride made me turn away.
The new ChatGPT voice has a MASSIVE vocal fry, like to ridiculous levels.
It did tone it down when I asked though, but a crazy default.
Moving mind in a moving body. Maybe the future of coding really is a philosophical debate with a dhjinn.. i like this a lot. bravo.
I actually can see this as advertisement USP. Some lone hiker going up a mountain through nature to a peak, while the city in the valley below far away becomes functional again because of what is said in that debate about .. walking for ideas - hunted down savanna mode similar to Darwin and all the other old thinkers.
lets expand on this idea.. augment reality overlay showing the resulting code in what you are looking at.. persistence is the ground.. model + business logic is the surroundings.. API is in the clouds..
Actually would pair better with functional programming.. as in the flow of the data - all those filter and mappers.. would fit well into the way you walk upon. But what is the center of your viewpoint and the center of attention, while you are moving. You would need to have a modal mode.. one where you move through the code structure by walking and/or looking and one - where the center is stationary, even while you move.
Oh wow, this is a great idea. Can I ask what your prompt looks like?
Me:
ChatGPT Live:
(I paste the link)
Me:
====
For the Socratic quiz I say:
I also have a Socratic quiz skill that I wrote for using in Claude Code or Codex to understand implementations/architecture etc:
https://pchalasani.github.io/claude-code-tools/plugins-detai...
thanks for sharing! saving this to refer back to this
also saved
To each their own, but for me - that'd be a heavy cognitive load. I ask Claude or ChatGPT to summarize HN comments and the article and provide any takeaways or wisdom nuggets. That's it.
His goal is to learn something. Your goal is to get information. Different scenarios.
Actually for me it’s more cognitive load when the AI just dumps everything at once, which is why I prefer to have it give me a nugget at a time, so I can take it in, pause and discuss before moving to the next one.
> You have to read the article
Nahh - pile in based on the title and what you assume the contents might be.
Did this while on a run for the past few days. The thing / interaction model I’ve been waiting for is here finally.
Thats an excellent idea. I've created https://gemini.google.com/gem/17xMogBqRSc2AtdCC-WRfObgHcceD0... and tested it with a couple conversations and it works great :)
Nice, it works for discussing concepts baked into the LLM but fails when you want it to read contents of an article or do web search. ChatGPT Live doesn’t have this limitation. Their Voice mode did have this limitation, but the Live mode released a couple weeks ago works exactly as you’d want, with link-following and web search
I can’t believe how bad Claude’s voice mode is. So much latency and it often tells me “this is rather technical” and would be better talking it out over text.
OpenAI has much less latency, can adapt to speak about any topic (though will tell you it put a code sample to look at later), and it just feels like brainstorming with a colleague.
If you haven’t tried the live voice version of ChatGPT recently, you should. I love to brainstorm ideas while on walks and it’s fabulous for that.
Ask it to summarize the conversation into a doc that will be waiting for you when you are back at your desk.
Is there a way to make the ChatGPT live experience less dumbed down?
It just wants to stay at a superficial level, while annoyingly mimicking human traits like “hmmms” and “umms”
I’ve found the better approach is to dictate into the text version and have it read the response. The responses are way better, so long as they don’t contain a table.
I like the podcast features still for getting an overview or orientation on a new topic where I know I need a big context dump before I can begin to do anything under my own speed.
I have tried the "walk and talk" pattern with AI, and it's okay, but I find that it can be kind of janky if the network connection isn't very good. And I find it very frustrating to use still in terms of how the interrupt feature works. For example, if I'm out hiking and there's even a little bit of wind, the voice mode of ChatGPT and Gemini both will think that I'm done talking and begin responding.
In fairness to the AI vendors, I imagine that this is a very difficult thing to get right. Humans have a very natural sense of flow in conversation, and knowing when a speaker is done talking and it's appropriate to begin responding.
notebook LM is still much much better at not hallucinating compared to claude and chatgpt.
Yes Claude voice refuses to actually read the article and relies on snippets it finds in web searches and often pretends to have read the article, and when pushed admits it really didn’t. ChatGPT Voice used to have this issue as well, but the new Live mode (released just a few weeks ago) fixed all that, and it actually pauses and reads the article and does real web searches during the conversation.
I never got into the podcasts side but it was good for synthesizing multiple documents in a way chatgpt or google's ai overview couldn't, but now I will just dump a couple PDFs into a new chat in Ollama and then ask it questions about the content. Is there another killer app for notebookLM out there?
>hack
>podcast slop
>letting the llm do it for you
there is a very good reason microsoft's ceo got repeatedly dunked on and it was because he literally couldn't stop babbling incoherently about having AI listen to things for him
i cannot imagine just sucking the joy out of life like this.
This may come as a surprise, but some people have to ingest large amounts of information for reasons other than producing joy.
what do you think will come of the people who have to ingest large amounts of information for reasons other than producing joy, if the AI is ingesting large amounts of information for them?
and anyway... the commenter is doing this for joy. so who, really, are you even talking about?
why be snarky? i agree, your AI is going to make the tedium in your job easier.
Probably nothing good, but that's not the complaint being raised!
> and anyway... the commenter is doing this for joy. so who, really, are you even talking about?
I don't see evidence of that?
you must be this guy
https://x.com/peregrinepulp/status/2077839461749338560?s=46&...
They do it for joy when they go on a walk, so perhaps what you find joyful are different than them, not sure why you care.
The limit of these things being at least partially due to the users’ imagination explains at least some of the terrible criticisms I see out there which do not stand up to scrutiny. Perfect example just now in the wild: https://x.com/peregrinepulp/status/2077839461749338560?s=46&...
anyway, your idea is sweet, likely because you are smart