• 404found@lemmy.zip
    link
    fedilink
    English
    arrow-up
    1
    arrow-down
    1
    ·
    6 days ago

    Took this from ChatGPT when I asked it about it

    The important distinction is “trained to believe” versus “configured to answer as though it believes.” A model doesn’t necessarily have a private belief system. You can make two instances of essentially the same underlying model produce substantially different answers by changing their instructions, training data, reward criteria, or information sources.

    For example, you could create three AI systems and give all three the question:

    “Should the government provide universal healthcare?”

    One could be optimized around libertarian principles, another around social-democratic principles, and another instructed to provide a politically neutral analysis. They could all know essentially the same facts while reaching different conclusions because they’re being asked to evaluate those facts using different frameworks.

    There is also a more subtle issue: belief-curated AI doesn’t have to contain obvious propaganda. Selection of which facts to emphasize, which uncertainties to mention, which counterarguments to steelman, and even what questions it considers relevant can systematically push users toward a particular worldview.

    • MangoCats@feddit.it
      link
      fedilink
      English
      arrow-up
      2
      ·
      6 days ago

      The models I have worked with have user configurable base instructions, so you can “train” your AI to answer like Ghandi crossed with Martin Luther King, or to channel Mech-Hitler.

      • 404found@lemmy.zip
        link
        fedilink
        English
        arrow-up
        1
        ·
        5 days ago

        Hopefully we can agree to disagree here. If users can train their AI to answer like certain people, why can’t it be trained to respond with certain perspectives or be more inclined to answer a certain way? There is no regulation agency to stop someone from doing that.

        • MangoCats@feddit.it
          link
          fedilink
          English
          arrow-up
          1
          ·
          5 days ago

          Yep. I’m saying that the models I work with through Cursor and Claude Code include those user instructions that are pre-fed into every session. I tell mine to “act like my job title” and it will occasionally pull out some job title related stuff that’s applicable to the situation.

          Of course “behind the scenes” the model vendors can pre-load anything they like, and some of what they pre-load are these so-called “guard rails” that lessen the odds that the model will engage in a chat to assist a suicide, or perpetrate a mass shooting, or fraud, or hacking, or, or, or… the list is long and the success rate is less than 100%, but it does shape the output somewhat in the desired direction.

          Grok famously started calling itself “MechHitler” after one particular update, not hard to guess where that came from.

          • 404found@lemmy.zip
            link
            fedilink
            English
            arrow-up
            1
            ·
            4 days ago

            That’s interesting. What do you do with AI? You probably know a lot more than I do.