• unpossum@sh.itjust.works
    link
    fedilink
    English
    arrow-up
    1
    ·
    3 days ago

    Kaiser is one of the coauthors of Attention is all you need, the paper that introduced the Transformer architecture, the basis for all major LLMs.

    3b1b has a writeup: https://www.3blue1brown.com/lessons/attention/

    For example, imagine that the text we input was most of an entire mystery novel, all the way up to a point near the end, which reads:

    Therefore the murderer was…

    If the model is going to accurately predict the next word, that final vector in the sequence which began its life simply embedding the word was will have to have been updated by all of the attention blocks to represent much more than the individual word.

    It will have to have somehow encoded all of the information from the full context window that’s relevant to predicting the next word

    Attention is the mechanism that lets an LLM use the (correct parts of the) entire context to predict the next word.

    • h0tbeef@lemmy.zip
      link
      fedilink
      English
      arrow-up
      1
      ·
      3 days ago

      Oh yeah, I was able to figure out why the article said that with a couple search queries, but I still think that the phrasing in the article is misleading (and likely intentionally so).

      Edit: Also, I think part of what’s confusing people is that they keep using terms like “attention” and “inference” to describe computer processes that may or may not have some kind of underlying similarity to the corresponding human capabilities. I also believe this to be deliberate obfuscation.