Tencent Releases and Open-Sources Tencent Hy4 Preview

(tencent.com)

92 points | by shenli3514 2 hours ago

9 comments

  • minimaxir 2 hours ago
    Hy4 apparently has ludicrous traction on OpenRouter already (https://openrouter.ai/tencent/hy4-preview), with trillions of tokens processed in a couple days: more than GLM 5.3 in a week. That said, it's relatively cheap with a 5% cache cost when everyone is still doing 10%/20% cache costs, so Hy4 may be more compelling.
    • martinald 1 hour ago
      I wrote about this a couple of weeks ago. It's actually often the biggest cost and it tends to be hidden away on most platforms!

      https://martinalderson.com/posts/watch-out-for-cache-read-co...

      Btw I still haven't came across any decent model that is <$0.01/MTok cache costs apart from deepseek thru their official API (even with the price increases).

      Seems like a bit of an opportunity for someone to take - drop cache read costs significantly.

      • dakolli 32 minutes ago
        That's because Deepseek invented the paradigm of prompt caching, they are the SOTA when it comes these techniques. Despite them open sourcing all their research, nobody beats them.

        edit: I do wish openrouter would let you sort providers by Cache Hit % and Cache cost. These are the only things that matter to me at this point when choosing a provider.

        • minimaxir 5 minutes ago
          You can click the table headers to sort Ascending/Descending.
    • Dinux 1 hour ago
      Which explains why almost none of my request go though
    • npn 1 hour ago
      [flagged]
    • cyanydeez 2 hours ago
      i'd be curious if openrouter is just being gamed by these publishers by paying for the exposure.

      wouldn't trust they dont do Capitalism like the rest of the AI field.

      • drob518 1 hour ago
        Of course they are. Of course they do. Nobody should be surprised by this.
      • tokai 1 hour ago
        >dont do Capitalism like the rest of the AI field

        Like lobbying the US president to harm their competitors?

        • realo 1 hour ago
          I would suggest "lobbying" is not the correct word to describe all the corruption going on in the current USA administration cesspool.
          • noir_lord 13 minutes ago
            lobbying/legalised bribery hard to say where one ends and another begins at times.
  • jorl17 1 hour ago
    I experimented with Hy3 for a project and was surprised with how good it was. I don't know if it's good for coding, but as a general purpose agentic model, it was only beaten by deepseek4-flash in our tests. It was so close to deepseek behaviour I kept thinking it must have been forked from it.
  • fastball 1 hour ago
    I wish model providers would stop committing chart crimes in their releases.

    - if you're gonna order the rest of the bar chart by rank, order your model accordingly.

    - if you're gonna highlight a winner in a table of benchmarks, don't highlight your entire model row in the table.

    Etc etc

    • mirekrusin 14 minutes ago
      Read websites through llm.
  • Zigurd 1 hour ago
    Is anyone here working on a problem for which current generation LLMs are inadequate, but that could possibly be solved by the next release of a first tier LLM?

    Or is it like bicycles? Unless your problem is named Tadej, you don't need a $13,000 bike.

    • comex 43 minutes ago
      My experience is that even Opus 5 still tends to write buggy or low-quality code and makes serious mistakes when analyzing code. It's a lot better than before but still not something I trust. I've had less experience with Fable since I can't use it at work; I hear it's a step up but still has its limits.

      For large tasks like a web browser or a compiler, even expensive swarms of frontier LLMs have not been shown capable of producing codebases that actually work. (Anthropic built a C compiler with Opus 4.6 but it lacked optimizations and apparently hit a complexity wall.)

      I also want to use LLMs for reverse engineering, but apparently it's pretty hit-or-miss, especially if you're forced to use open-source models to avoid restrictions.

    • lopatin 31 minutes ago
      I asked a current generation LLM to make me $1k a week and it hasn't so far.
    • tekacs 12 minutes ago
      Yes, lots – I think that folks will hopefully discover more of these as they scale up their ambition, now that LLMs make a lot of previously difficult things far easier.
    • ezst 33 minutes ago
      I saw a laptop earlier in the train that I asked ChatGPT, Claude and Gemini what it was, providing a brand, screen size and ports description. Gemini could never figure it out, Claude and ChatGPT eventually did, after multiple rounds of indirection, giving completely wrong answers (there was a perfect match for the problem statement, they all explored alternatives first). LLMs are (probably) amazing at things I don't care about, and still suck at the mundane stuff you would have the marketing tell you they excel at.
    • RGS1811 1 hour ago
      For me personally, the answer is no. Fable is adequate to do basically anything I want to do. My perspective, broadly speaking, is that we've saturated most of the benchmarks because we've largely saturated our capacity to verify models' work at scale. What's left is context-bound verification, i.e. the problem of ensuring that output matches intent and ambiguities in prompting were resolved correctly. Further advances in autonomy do not make that latter verification problem easier. If anything they make it harder as the output per task becomes more complex and therefore more taxing for a human to verify.

      The solution to that (to my mind) would be not a better model but a basic shift in architecture beyond the current paradigm and into a setup where agents have durable, plastic memories and undergo contextual individuation over time. But at that point agents start to become quasi-persons and not tools.

    • er4hn 36 minutes ago
      I was given a picture cube, which is like a Rubik's cube but every side is a unique picture. It came scrambled and I don't have an original reference image. I like to take videos of it and give it to llms to solve. I call it my agi test because it hasn't been solved yet
    • spacebanana7 40 minutes ago
      I want to be able to generate my own Simlilirian movie by dumping the content of a book into an LLM.

      Both animated and live action results would be acceptable.

      Unfortunately most existing LLMs lack the capability to maintain context across tens of thousands of frames.

      • Demiurge 8 minutes ago
        That sounds like an interesting challenge. Have you seriously considered solving it? Because in about 10 seconds I came up with a process that should work, provided enough compute power. Simply model the traditional film making process by starting with a script, character stories. Design your world, then design the storyboard, and all the scenes. Create a list of all the visual elements that need to be replicated between all the scenes. Then you have to built prompts and reference art of the objects, faces, people. Make sure to do multiple takes of each scene, and have the vLLM critique and analyze the performances and technicalities. Should work?

        I think, also, like in the traditional film makers career, this process should be built iteratively, start with a fast food commercial, then do a music video, then you can probably do a short film. Continue to improve the process, and one day I’m sure the LLM film studio can make you any movie you want, provided you have enough tokens.

        • andybak 1 minute ago
          I'm getting a Poe's Law feeling. I'm genuinely unsure about whether this post is a stone cold parody or not. I think I need to turn off the internet and go to bed.
    • _factor 1 hour ago
      Hardware debugging and firmware details lead to thinking/testing loops on all but the frontier here.
    • dakolli 35 minutes ago
      I get buy with very cheap models and actually using my brain, you don't need these SOTA models. China will definitely win this AI 'war'
      • kennywinker 5 minutes ago
        If the models stay open, it seems like everybody but anthropic/openai wins. i literally can’t see a downside. We can post-train the models to know about tienanmen square.
    • tokai 1 hour ago
      A spanish rock solved that problem for free.
    • poincareball 1 hour ago
      [dead]
  • XCSme 39 minutes ago
    I tried benchmarking it, but it keeps timing out/rate limiting, so the current provider(s) are unusable.
  • vcryan 1 hour ago
    I used Hy3 quite a bit for the type of tasks it was suited for. Excited about this. My one concern over Hy3 was speed. In theory, it could be served much faster as a smaller model but it was relatively slow everywhere I could get it (including from Tencent directly) but also several other inference providers.
    • Topfi 1 hour ago
      In my evals, I saw an unprecedented jump between preview and final release on Hy3, from unusable to competitive. Did you see similar in preview vs release version?
      • vcryan 1 hour ago
        Oh yes! I forgot about that. Yes, you can see this in benchmarks about hy3 preview and hy3 release still today because they measured them separately - it was significant.
  • usernomdeguerre 1 hour ago
    is it just me or are the bar charts in the blog post strange? Higher numbers don't seem to correspond correctly to their actual height?
    • pixelesque 30 minutes ago
      Looks okay to me.

      The first column has both the Hy4 and Hy3 scores overlaid on one another (Hy4 is darker blue and the taller one), with both scores written below the top of the respective bar - maybe you're seeing that?

    • alanfranz 1 hour ago
      Probably AI generated.

      But, what bars are clearly off? I couldn't spot any.

    • feynmanquest 1 hour ago
      Noticed that as well
  • onesandofgrain 1 hour ago
    Go go China!

    EDIT: I don't know why I'm being downvoted by AI bots.

    • yogthos 47 minutes ago
      Seriously, without China we'd just have two parasitic companies hoarding this tech and deciding whom and how is allowed to use it.
  • petcat 49 minutes ago
    > Tencent has released and open-sourced Tencent Hy4 preview, a next-generation large language model with 770B total parameters and 49B active parameters, and a context window exceeding 1M tokens.

    There are no open source models, at least not useful ones (yet) [0]. Open weight is not the same as open source. The current "open weight" models are just opaque binary blobs you can run on your own computer instead of through a web API.

    [0] https://allenai.org/

    Imagine thinking that running a Photoshop binary on your own computer instead of through a SaaS web app means that it's "open source". Of course you think that's ridiculous.

    • mirekrusin 12 minutes ago
      You can open source dataset without all the details how it was assembled.

      Models are lossy compressed datasets you can pick up and amend (fine tune / continue training / alter) according to license they were released under.

      Hy4 is released under OSI approved Apache License 2.0.