Technology

59404 readers

2077 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related content.
Be excellent to each another!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, to ask if your bot can be added please contact us.
Check for duplicates before posting, duplicates may be removed

Approved Bots

founded 1 year ago

MODERATORS

490

In Cringe Video, OpenAI CTO Says She Doesn’t Know Where Sora’s Training Data Came From (futurism.com)

submitted 8 months ago by ylai@lemmy.ml to c/technology@lemmy.world

185 comments fedilink hide all child comments

you are viewing a single comment's thread
view the rest of the comments

[–] GiveMemes@jlai.lu -1 points 8 months ago (1 children)

Sorry, I know reading the whole article is hard:

The complaint cites several examples when a chatbot provided users with near-verbatim excerpts from Times articles that would otherwise require a paid subscription to view.

[–] A_Very_Big_Fan@lemmy.world 2 points 8 months ago* (last edited 8 months ago)

Yeah lmao after like 20 paragraphs of nothing, it wasn't hard to believe you didn't know what you were talking about. But I looked at the complaint itself out of curiosity, and it's flimsy and misleading.

The first issue is 100% of the allegedly paywalled text from all 4 articles mentioned in the complaint can be read by non-paying customers for free outside of the paywall. You can't read the whole article, but you can get far enough to read all 4 quotes mentioned in the complaint yourself. The links to each article are in the complaint if you don't believe me. They have nothing to show they bypassed a paywall or that it was trained on unlicensed content.

The second issue is the third exhibit claims it will bypass paywalls when asked. This is demonstrably false because for one, the article they asked it for isn't paywalled, and for two, using their exact prompts word for word doesn't work if you try it yourself.

Two of the four exhibits don't even have screenshots, so there's no evidence it happened in the first place, but more importantly they don't (and apparently won't when asked) disclose what lengths they had to go to in order to get that output. For all we know they gave it 90% of the words and told it to fill in the gaps.