Comment

Comment on Bill proposed to outlaw downloading Chinese AI models.

There are several “good” LLMs trained on open datasets like FineWeb, LAION, DataComp, etc. They are still “ethically dubious”, but at least they can be downloaded, analyzed, filtered, and so on. Unfortunately businesses are keeping datasets and training code as a competitive advantage, even "Open"AI stopped publishing them when they saw an opportunity to make money.

What is the concern with only having weights? It’s not abritrary code exectution

Unless one plugs it into an agent… which is kind of the use we expect right now.

Accessing the web, or even web searches, is already equivalent to arbitrary code execution: an LLM could decide to, for example, summarize and compress some context full of trade secrets, then proceed to “search” for it, sending it to wherever it has access to.

Agents can also be allowed to run local commands… again a use we kind of want now (“hey Google, open my alarms” on a smartphone).

source

Sort:hotnew top

p03locke@lemmy.dbzer0.com ⁨1⁩ ⁨year⁩ ago

There are several “good” LLMs trained on open datasets like FineWeb, LAION, DataComp, etc.

Then use those as training data. You’re too caught up on this exacting definition of open source that you’ll completely ignore the benefits of what this model could provide.

an LLM could decide to, for example, summarize and compress some context full of trade secrets, then proceed to “search” for it, sending it to wherever it has access to.

That’s not how LLMs work, and you know it. A model of weights is not a lossless compression algorithm.

Also, if you’re giving an LLM free reign to all of your session tokens and security passwords, that’s on you.

source
- jarfil@beehaw.org ⁨1⁩ ⁨year⁩ ago
  
  That’s not how LLMs work, and you know it. A model of weights is not a lossless compression algorithm.
  
  piratewires.com/…/compression-prompts-gpt-hidden-…
  
  if you’re giving an LLM free reign to all of your session tokens and security passwords, that’s on you.
  
  There are more trade secrets than session tokens and security passwords. People want AI agents to summarize their local knowledge base and documents, then expand it with updated web searches. No passwords needed when the LLM can order the data to be exfiltrated directly.
  
  source
teawrecks@sopuli.xyz ⁨1⁩ ⁨year⁩ ago
Those security concerns seem completely unrelated to the model, though. You can have a completely open source model that fits all those requirements, and still give it too much unfettered access to important resources with no way of actually knowing what it will do until it tries.

source
- jarfil@beehaw.org ⁨1⁩ ⁨year⁩ ago
  While unfettered access is bad in general, DeepSeek takes it a step farther: the Mixture of Experts approach in order to reduce computational load, is great when you know exactly what “Experts” it’s using, not so great when there is no way to check whether some of those “Experts” might be focused on extracting intelligence under specific circumstances.
  
  source
  - teawrecks@sopuli.xyz ⁨1⁩ ⁨year⁩ ago
    I agree that you can’t know if the AI has been deliberately trained to act nefarious given the right circumstances. But I maintain that it’s (currently) impossible to know if any AI had been inadvertently trained to do the same. So the security implications are no different. If you’ve given an AI the ability to exfiltrating data without any oversight, you’ve already messed up, no matter whether you’re using a single AI you trained yourself, a black box full of experts, or deepseek directly.
    
    But all this is about whether merely sharing weights is “open source”, and you’ve convinced me that it’s not. There needs to be a classification, similar to “source available”; this would be like “weights available”.
    
    source