跳到正文

#语音

今日 0 条
9月29日周二
  1. Tesla34

    在你的 Tesla 里使用 Grok @Bot

    引用Mike P@mikepat711

    I dramatically underestimated the value of Grok Bot in my Tesla. Can you technically do everything you’re able to do with Grok Bot in the Tesla just by launching the app on your phone and going voice mode? Sure. But the magic isn’t in the raw capability itself, it’s in the combination of raw capability and total seamlessness. I often have fleeting thoughts while driving around. Some of these things aren’t significant enough to make me fumble around with my phone while supervising FSD, but when I literally don’t have to lift a finger and can just “Hey Grok” anything computer related into reality, I find myself taking advantage daily. And the little things add up. I’m on my way to work and remembered that I have a few documents on my home computer that I am going to need for work. So I just asked “Hey Grok, can you grab those 4 pdf files from my Mac downloads folder and move them into X folder on my one drive?” Now they’ll be where I need them to be when I get to my desk, and I won’t have to re-download them on my work machine. I have to work on a presentation today for a talk in a few weeks. “Hey grok, go into my sales enablement folder and have Claude Code build me a deck focused on X product for Y industry.” Claude already has my full job context, styling preferences etc and can build a nearly finished deck with a prompt. Now when I get to my desk, instead of building from scratch, I can review and prompt revisions before beginning to dig through my email backlog. Feels insane.

9月24日周四
  1. Google Gemini 博客71

    Google 发布 Gemini 3.8 Live with Live Avatar

    Google 发布 Gemini 3.8 Live with Live Avatar,将低延迟流式视频与实时对话能力结合,为对话模型带来近实时的视觉形象。该功能支持精确唇形同步、自然表情与流畅轮次切换,并可在对话中异步调用工具、在后台获取数据。它支持 97 种语言的原生语音到语音同步,输出均带 SynthID 水印,现已在 Gemini Enterprise 提供,自定义形象仅限企业白名单申请。

    推荐理由:Gemini 3.8 Live 新增实时视频形象能力,读者可了解多模态对话在企业场景中的落地形态与开放范围。

  2. Higgsfield Blog18

    2026 年 15 款最佳 AI 内容创作平台盘点

    Higgsfield Blog 盘点 2026 年 15 款面向内容创作者的 AI 平台,覆盖策划、生成、剪辑与发布全流程。Higgsfield 以 30+ 模型整合图像、视频与广告制作,起价 $15/mo;Canva 主打日常设计,Descript 与 OpusClip 侧重长视频二次剪辑,HeyGen 与 Synthesia 专注数字人出镜视频。

9月23日周三
  1. Google Gemini 博客72

    Google 发布 Gemini 3.8 Flash TTS 与 Flash-Lite TTS

    Google 发布 Gemini 3.8 Flash TTS 和 Gemini 3.8 Flash-Lite TTS 两款文本转语音模型,前者面向角色设计与深度创作控制,后者面向高并发、低成本的配音与语音智能体场景。

    推荐理由:两款 TTS 模型把语音生成从固定音色扩展到自然语言定制与逐句导演,并给出基准排名和开放入口。

9月19日周六
9月17日周四
  1. SpaceXAI58

    xAI 的 Grok Voice 已在 fal 上线,可用于构建低延迟、能解决真实客户问题的智能体。该语音模型响应时间为 0.70 秒,可在句子结束前完成工具调用,支持带词级时间戳的转写、25 种以上语言的 30 多种音色文本转语音,并可用两分钟音频进行声音克隆。

    引用fal@fal

    Grok Voice is live on fal. The latest speech model from @SpaceXAI that answers in 0.70 seconds and finishes its tool calls before the sentence ends. Transcription with word-level timestamps, text to speech in 30+ voices across 25+ languages, and cloning from two minutes of audio. Every voice in this video is from Grok Voice.

9月16日周三
9月9日周三
9月2日周三
9月1日周二
  1. Meta AI Research · Muse74

    Meta 发布 Muse Voice Transcribe 实时语音感知模型

    Meta Superintelligence Labs 发布 Muse Voice Transcribe,这是其首款实时音频感知模型,支持流式 ASR、20 人以上说话人分离和端点检测。模型原生支持任意中英混说,训练覆盖 70 多种语言,其中 25 种经过充分验证,并支持超过一小时的长音频输入。官方称其在 Artificial Analysis 流式语音转文字和公开说话人分离基准上排名第一。

    推荐理由:Meta 公布实时语音感知模型的技术细节与基准排名,读者可了解流式 ASR 与说话人分离的实现路径。

8月3日周一
7月30日周四
  1. Google Labs 博客58

    Google 发布音乐生成模型 Lyria 3.5,上线 Flow Music

    Google 发布新一代音乐生成模型 Lyria 3.5,并已在 Google Flow Music 中上线。官方称其在音乐性、歌词、人声和创作控制四方面均有提升:可生成更丰富自然的旋律结构,歌词质量与提示词遵循度、结构感知更好,人声更具表现力和情感且发音改善,同时更易控制输出的节奏与时长。

6月29日周一
  1. 十字路口Crossing34

    越伴动力创始人世博谈陪伴机器人「小伴」:不说人话、95% 柔软材质、0.4 秒交互延迟

    越伴动力发布陪伴机器人「小伴」,它不说人话,而是发出类似"外星语"的声音,能撒娇、委屈和拒绝。产品全身 95% 为柔软材质,采用端侧 1.7B 快脑与 7B 慢脑分工,将交互延迟压到 0.4 秒以内,并配备云端超长程记忆推动性格动态演化。创始人世博提出陪伴不是讨好、生命力不是可爱、少就是多三条产品判断。

5月28日周四
10月27日周一
  1. 十字路口Crossing26

    未来智能马啸谈 AI 耳机:亿元级融资后如何挑战硬件“不可能三角”

    未来智能创始人兼 CEO 马啸在完成蚂蚁、启明等投资的亿元级融资后,分享了 AI 耳机的产品取舍逻辑:在续航、重量与处理性能的“不可能三角”中,将通话续航做到 9-10 小时,高于市面常见的 5-6 小时。其讯飞 AI 耳机已实现持续盈利,并稳坐各大电商平台 AI 耳机销量榜首。

9月22日周一
  1. 十字路口Crossing30

    前作业盒子创始人刘夜再创业:VisionFlow 推出 AI 口语产品 Talkit

    前作业盒子创始人刘夜创办 VisionFlow,获约 1000 万美元种子轮融资,出资方包括李想、阿里巴巴合伙人曾鸣及语嫣,并推出首款产品 Talkit——一个为口语练习打造的 AI x 3D 虚拟世界。团队自研“世界生成引擎(Gen World Engine)”,称实习生一个月可生成 1000 个 3D 虚拟人。刘夜认为 AI 大模型让诞生于 1980 年的 TBLT 语言学习理论真正迎来春天。

5月14日周二
  1. Sam Altman Blog88

    OpenAI 发布 GPT-4o,免费向 ChatGPT 用户开放并推出语音视频模式

    OpenAI 发布 GPT-4o,Sam Altman 表示已在 ChatGPT 中免费提供这一模型,且不含广告。他称新的语音与视频模式是自己用过最好的计算机界面,达到接近人类的响应速度和表现力,并提到后续会加入可选个性化、访问用户信息以及代为执行操作等能力。

    推荐理由:OpenAI 官方说明 GPT-4o 免费开放与语音视频模式,读者可了解其能力定位与商业化取舍。

4月6日周六