points by d4rkp4ttern 1 week ago

When notebookLM was new, it was interesting to listen to the podcasts. Then the novelty wore off, and I wanted something where I can interact with the podcasters but it was janky as hell.

My current “audio-learning” hack is ChatGPT Live which has become shockingly good after being awful compared to Claude Voice (Let’s not even talk about Gemini voice which is still bad).

I go on a walk and dump a paper or article link in the chat, and ask chatGPT Live to walk me through the content in small nuggets, so I can discuss them interactively. For deeper topics I have it quiz me Socratic style so I’m not just passively listening, and actually thinking through problems or ideas.

siquick 1 week ago

There is actually an interrupt mode in NotebookLM now.

Overall I've found it the best AI product Google have. Only complaint I have about it is the hyper positive US corporate accents get pretty annoying pretty quickly.

The realtime voice in ChatGPT is excellent, the newer model is a big step up too.

  • fumeux_fume 1 week ago

    I wonder how much of the annoying aspects of the podcasters can be controlled by the custom user instructions. I hate sitting through the patter of how "we're not political, just reporting the views of the sources."

    • idbnstra 1 week ago

      "We have some REALLY interesting sources that we're discussing today"

  • toddmorey 1 week ago

    Yeah I think the podcasts were just a clever hack until AI is fast enough to support realtime. It’s practically there now, so a fake podcast is a funny relic and hard to sit through.

  • d4rkp4ttern 1 week ago

    Too much friction in notebookLM - several minutes to make a podcast, and they disappear after a while, and the interrupt feature is super-janky.

    • AbstractH24 1 week ago

      > Too much friction in notebookLM

      This was my main problem with it. Loved it otherwise, but the lack of an API or other automation so I could dump stuff in it and have a podcast ready to download and listen to on my subway ride made me turn away.

  • theshrike79 1 week ago

    The new ChatGPT voice has a MASSIVE vocal fry, like to ridiculous levels.

    It did tone it down when I asked though, but a crazy default.

21asdffdsa12 1 week ago

Moving mind in a moving body. Maybe the future of coding really is a philosophical debate with a dhjinn.. i like this a lot. bravo.

I actually can see this as advertisement USP. Some lone hiker going up a mountain through nature to a peak, while the city in the valley below far away becomes functional again because of what is said in that debate about .. walking for ideas - hunted down savanna mode similar to Darwin and all the other old thinkers.

  • 21asdffdsa12 1 week ago

    lets expand on this idea.. augment reality overlay showing the resulting code in what you are looking at.. persistence is the ground.. model + business logic is the surroundings.. API is in the clouds..

    • 21asdffdsa12 1 week ago

      Actually would pair better with functional programming.. as in the flow of the data - all those filter and mappers.. would fit well into the way you walk upon. But what is the center of your viewpoint and the center of attention, while you are moving. You would need to have a modal mode.. one where you move through the code structure by walking and/or looking and one - where the center is stationary, even while you move.

citiguy 1 week ago

Oh wow, this is a great idea. Can I ask what your prompt looks like?

  • d4rkp4ttern 1 week ago

    Me:

        Ok so I'm going on a walk. I'll dump a link to a Hacker News   
        discussion about an article. 
        You have to read the article and the discussion and walk me thru   
        all the interesting details, nugget by nugget, and move on when 
        I'm ready for the next piece.
    

    ChatGPT Live:

        ok, Great show me the link, I'm waiting.
    

    (I paste the link)

    Me:

        Ok I pasted it. Now go.
    

    ====

    For the Socratic quiz I say:

        I want to understand this more deeply. So instead of you just telling me
        everything, lay out the problem and a question for me to think about, and 
        I'll try to answer. Even if I answer wrong, you should resist giving me the
        answer, and instead keep digging with more questions, so that I eventually 
        arrive at the answer myself.
    

    I also have a Socratic quiz skill that I wrote for using in Claude Code or Codex to understand implementations/architecture etc:

    https://pchalasani.github.io/claude-code-tools/plugins-detai...

    • singhkays 1 week ago

      thanks for sharing! saving this to refer back to this

    • mandeepj 1 week ago

      To each their own, but for me - that'd be a heavy cognitive load. I ask Claude or ChatGPT to summarize HN comments and the article and provide any takeaways or wisdom nuggets. That's it.

      • john_minsk 1 week ago

        His goal is to learn something. Your goal is to get information. Different scenarios.

      • d4rkp4ttern 1 week ago

        Actually for me it’s more cognitive load when the AI just dumps everything at once, which is why I prefer to have it give me a nugget at a time, so I can take it in, pause and discuss before moving to the next one.

    • blitzar 1 week ago

      > You have to read the article

      Nahh - pile in based on the title and what you assume the contents might be.

    • _boffin_ 1 week ago

      Did this while on a run for the past few days. The thing / interaction model I’ve been waiting for is here finally.

    • windenntw 1 week ago

      Thats an excellent idea. I've created https://gemini.google.com/gem/17xMogBqRSc2AtdCC-WRfObgHcceD0... and tested it with a couple conversations and it works great :)

      • d4rkp4ttern 1 week ago

        Nice, it works for discussing concepts baked into the LLM but fails when you want it to read contents of an article or do web search. ChatGPT Live doesn’t have this limitation. Their Voice mode did have this limitation, but the Live mode released a couple weeks ago works exactly as you’d want, with link-following and web search

toddmorey 1 week ago

I can’t believe how bad Claude’s voice mode is. So much latency and it often tells me “this is rather technical” and would be better talking it out over text.

OpenAI has much less latency, can adapt to speak about any topic (though will tell you it put a code sample to look at later), and it just feels like brainstorming with a colleague.

If you haven’t tried the live voice version of ChatGPT recently, you should. I love to brainstorm ideas while on walks and it’s fabulous for that.

Ask it to summarize the conversation into a doc that will be waiting for you when you are back at your desk.

  • scarab92 1 week ago

    Is there a way to make the ChatGPT live experience less dumbed down?

    It just wants to stay at a superficial level, while annoyingly mimicking human traits like “hmmms” and “umms”

    I’ve found the better approach is to dictate into the text version and have it read the response. The responses are way better, so long as they don’t contain a table.

_doctor_love 1 week ago

I like the podcast features still for getting an overview or orientation on a new topic where I know I need a big context dump before I can begin to do anything under my own speed.

I have tried the "walk and talk" pattern with AI, and it's okay, but I find that it can be kind of janky if the network connection isn't very good. And I find it very frustrating to use still in terms of how the interrupt feature works. For example, if I'm out hiking and there's even a little bit of wind, the voice mode of ChatGPT and Gemini both will think that I'm done talking and begin responding.

In fairness to the AI vendors, I imagine that this is a very difficult thing to get right. Humans have a very natural sense of flow in conversation, and knowing when a speaker is done talking and it's appropriate to begin responding.

msh 1 week ago

notebook LM is still much much better at not hallucinating compared to claude and chatgpt.

  • d4rkp4ttern 1 week ago

    Yes Claude voice refuses to actually read the article and relies on snippets it finds in web searches and often pretends to have read the article, and when pushed admits it really didn’t. ChatGPT Voice used to have this issue as well, but the new Live mode (released just a few weeks ago) fixed all that, and it actually pauses and reads the article and does real web searches during the conversation.

ctkhn 1 week ago

I never got into the podcasts side but it was good for synthesizing multiple documents in a way chatgpt or google's ai overview couldn't, but now I will just dump a couple PDFs into a new chat in Ollama and then ask it questions about the content. Is there another killer app for notebookLM out there?

lardosaurusrex 1 week ago

>hack

>podcast slop

>letting the llm do it for you

there is a very good reason microsoft's ceo got repeatedly dunked on and it was because he literally couldn't stop babbling incoherently about having AI listen to things for him

i cannot imagine just sucking the joy out of life like this.

  • estearum 1 week ago

    This may come as a surprise, but some people have to ingest large amounts of information for reasons other than producing joy.

    • doctorpangloss 1 week ago

      what do you think will come of the people who have to ingest large amounts of information for reasons other than producing joy, if the AI is ingesting large amounts of information for them?

      and anyway... the commenter is doing this for joy. so who, really, are you even talking about?

      why be snarky? i agree, your AI is going to make the tedium in your job easier.

      • estearum 1 week ago

        Probably nothing good, but that's not the complaint being raised!

        > and anyway... the commenter is doing this for joy. so who, really, are you even talking about?

        I don't see evidence of that?

  • satvikpendem 1 week ago

    They do it for joy when they go on a walk, so perhaps what you find joyful are different than them, not sure why you care.

pmarreck 1 week ago

The limit of these things being at least partially due to the users’ imagination explains at least some of the terrible criticisms I see out there which do not stand up to scrutiny. Perfect example just now in the wild: https://x.com/peregrinepulp/status/2077839461749338560?s=46&...

anyway, your idea is sweet, likely because you are smart