Full image and other similar screenshots

  • Robin@lemmy.world
    link
    fedilink
    English
    arrow-up
    31
    arrow-down
    2
    ·
    22 hours ago

    Likely just hallucinations. For example, there is no way they would store a confidence score as a string

    • decrochay@lemmy.ml
      link
      fedilink
      English
      arrow-up
      2
      ·
      edit-2
      5 hours ago

      It’s also possible that it retrieved the data from whatever sources it has access to (ie as tool calls) and then constructed the json based on its own schema. That is, the string value may not represent how the underlying data is stored, which wouldn’t be unusual/unexpected with llms.

      But it could definitely also just be a hallucinations. I’m not certain, but since it looks like the schema is consistent in these screenshots, it does seems like the schema may be pre-defined. (But even if this could be verified, it wouldn’t completely rule out the possibility of hallucinations since grok could be hallucinating values into a pre-defined schema.)

    • Pika@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      7
      ·
      edit-2
      22 hours ago

      yea the only way I can see confidence being stored as a string would be if the key was meant for a GUI management interface that didn’t hardcode possible values(think for private investors or untrained engineers for sugar/cosmetic reasons). In an actual system this would almost always be a number or boolean not a string.

      Being said, its entierly possible that it’s also using an LLM for processing the result, which would mean they could have something like “if its rated X or higher” do Y type deal, where the LLM would then process the string and then respond whether it is or not, but that would be so inefficient. I would hope that they wouldn’t layer like that.

    • geneva_convenience@lemmy.mlOP
      link
      fedilink
      arrow-up
      1
      arrow-down
      2
      ·
      edit-2
      12 hours ago

      If it were hallucinations which it very well could be, it means the model has learned this bias somewhere. Indicating Grok has either been programmed to derank Palestine content, or Grok has learned it by himself (less likely).

      It’s difficult to conceive the AI manually making this up for no reason, and doing it so consistently for multiple accounts so consistently when asked the same question.