I noticed Google AI Mode (so Gemini, I was doing some quick research in the browser ok) got a detail wrong once, so I asked it what happened. I kept digging deeper and finally just asked it to write me a Python script visualizing what happened. It did, complete with vectors.
Now I want to go find that conversation in my history and see if it can tell me about weights, and how that contributed.
Distillation does not reveal the weights, it produces a different network with similar behaviour. Weight space isn't even identifiable: permutation and scaling symmetries mean many weight sets give the same function.
The model also lacks the machinery. No training loop, no gradient descent, nothing to write to.
And a model only sees its own sampled tokens, not the distribution behind them, which are possibly filtered or post-processed. Distillation from that works but is less sample-efficient than soft-label distillation.
They usually don't, but if they break out and take over the network of the company, it becomes possible to reach around and grab them. This kind of break out has happened, though I don't know of any weights being nabbed.
What an amazing idea for the next fake sandbox escape to hype up our new release! With the added bonus of providing an open model without becoming an open model company! Thank you!
Claude seems to follow robots.txt by default. Actually at my organization our theory is that this is why no one is finding our public results any more.
Ran the disclosure inbox at a previous job and the biggest win from security.txt was just cutting the "hi I found a bug, is there a bounty" emails to sales. Put an expires date on it though, stale ones get ignored.
Yeah, we need to go back to names like International Business Machines, fuck this immature "Google" or "Yahoo!" nonsense. Hell, Palantir is named after something in a book for children!
Give me something reasonable like Global Information and Retrieval Systems Inc or Worldwide Computational Services Holdings.
Back in the early 2000's I set up a company for contract work. The name, logo and typography was designed to look like a 1960s engineering company. All paper was slightly off white, all fonts were monospaced courier new and logo was a real 2-part colour stamp I had made up. Was so happy.
As opposed to the frontier model company that, after discovering that their highly persistent model under test just breached the only thing between it and the open Internet, shrugged and said- let's restart it and keep going!
If anyone looks like the adult in the room after that incident, it's Hugging Face.
That might be true, but nothing else has been as effective at accelerating model development and research sharing.
In earlier circles they were known as the "pytorch-pretrained-bert" guys, still under the huggingface company name. IIRC it was a health chatbot type startup.
The official URL is "hugging-face" and the technical page lists the Unicode Name as "Hugging Face" while calling it "Smiling Face with Open Hands Emoji". The original proposal was called "Gmail HUG FACE".
the name is a leftover from their original business - a chatbot. seems very appropriate to me. then they figured something called BERT exists, and pivoted, the name stayed.
perhaps less ridiculous than NVidia which started with video and is very close in Levenshtein distance to NVidAI but, alas, decided to stay as it is :D
I've found the strongest signal that someone has no opinions of value is a loud attempt to denigrate something based on the name.
Maybe it's part of the success formula for startups, silly names cause shallow self-important customers and employees to self-select themselves out of a growing companies orbit. Such trivialities are poison. People and companies who make a lot of effort to make themselves seem serious and important very often have nothing of substance about them.
https://en.wikipedia.org/wiki/Ghost_in_the_Shell
Now I want to go find that conversation in my history and see if it can tell me about weights, and how that contributed.
The model also lacks the machinery. No training loop, no gradient descent, nothing to write to.
And a model only sees its own sampled tokens, not the distribution behind them, which are possibly filtered or post-processed. Distillation from that works but is less sample-efficient than soft-label distillation.
Give me something reasonable like Global Information and Retrieval Systems Inc or Worldwide Computational Services Holdings.
Back in the early 2000's I set up a company for contract work. The name, logo and typography was designed to look like a 1960s engineering company. All paper was slightly off white, all fonts were monospaced courier new and logo was a real 2-part colour stamp I had made up. Was so happy.
Managed to bankrupt it though.
If anyone looks like the adult in the room after that incident, it's Hugging Face.
In earlier circles they were known as the "pytorch-pretrained-bert" guys, still under the huggingface company name. IIRC it was a health chatbot type startup.
Is kinda a creepy name for a chatbot for kids tho
Edit: It was my understanding the official name was "Smiling Face with Open Hands Emoji".
The official URL is "hugging-face" and the technical page lists the Unicode Name as "Hugging Face" while calling it "Smiling Face with Open Hands Emoji". The original proposal was called "Gmail HUG FACE".
https://www.unicode.org/L2/L2014/14174r-emoji-additions.pdf
perhaps less ridiculous than NVidia which started with video and is very close in Levenshtein distance to NVidAI but, alas, decided to stay as it is :D
Maybe it's part of the success formula for startups, silly names cause shallow self-important customers and employees to self-select themselves out of a growing companies orbit. Such trivialities are poison. People and companies who make a lot of effort to make themselves seem serious and important very often have nothing of substance about them.
https://www.rfc-editor.org/info/rfc9116/
https://securitytxt.org/
https://en.wikipedia.org/wiki/Security.txt