Conversation

Ben Lubar (any pronouns)

determining whether you're arguing with an LLM is quite simple: ask yourself "is this person smart enough to even know what an em dash is" and if the answer is no and they're using em dashes, well...

1
0
0

Not sure why we need "text watermarking" according to the EU when there are a billion different tells that you're reading something written by the bullshit autocomplete machine.

2
0
0

https://steamcommunity.com/discussions/forum/10/587308261888267961/?ctp=5

"Why doesn't Steam simply tell video games where to put their save files?"

IT DOES. That's literally the purpose of the deprecated function I linked to.

They deprecated it because they figured out reality doesn't work like that.

0
0
0

@ben it’s so that even absent of such tells (which are not intrinsic of LLM generations, just an artifact of how current ones work), there should still be a way to determine whether something is LLM generated or not

1
0
1

@charlotte ok but why not just ban LLM chatbots then

they're misinformation generators

the only service they provide is confidently phrased misinformation

2
0
0

@ben banning LLMs in one geographical region will not make LLM text disappear

0
0
0

@charlotte or is the EU's plan to get the chatbots to voluntarily discontinue their service in the EU because "watermarking" randomly generated text in a way that survives even the simplest of edits is nigh-impossible

and then it backfired when anthropic went "we'll just bullshit them and claim we did it"

2
0
0

@charlotte I guess you could have some kind of algorithm that always picks a specific word in the probability list under certain conditions but that requires having the entire model and running it again to determine whether the text might match

that's not a watermark, that's a heuristic

1
0
0

@ben it’s possible to watermark text that survives some editing (copying and pasting)

0
0
0

@ben it’s also possible to bias the LLM word choice in a way where it can be detected statistically with something much smaller than an LLM

1
0
0

@charlotte inferencing isn't terribly expensive; it's the training that screws up the environment

my issue with the LLM's choices being where the watermark exists is that it's completely opaque to anyone who isn't the company that owns the LLM

so the LLM company would charge people money to say whether it's likely that they generated a specific piece of misinformation

1
0
0

Charlotte lotteheartplural/Cinny cinny_heart_plural thetadelta ursaminor treblesand

Edited 12 days ago

@ben the EU did think of that

the watermark ought to be interoperable between vendors as far as is feasible, this will become a requirement next year

there is also an associated code of practice that would be contractually binding to AI companies. openai, anthropic, google, microsoft, meta, and mistral have signed it, along with over 150 other companies. one stipulation is that the detection needs to be

  • always free for specific groups (including researchers, fact checkers, journalists, law enforcement)
  • free if the company has >1mil MAU
  • free if it’s not bulk use
  • and the fee needs to be proportional to the operational cost and be significant for the company
0
0
1