Hacker Newsnew | past | comments | ask | show | jobs | submit | jonathanstrange's commentslogin

Governments don't necessarily have the connection between your ID and what you do online, though, and some governments are known to massively buy publicly available data from data brokers to circumvent existing laws. By "some governments" I mean the US government, by the way. That's not a conspiracy either, it's well-documented.

The technology was first released into the public 4 years ago. Now it's proving major theorems in research mathematics. It doesn't take a genius to realize that this might indicate an extinction-level shift over the next few hundred years.

Academic work is based on worldwide sharing, the sharing is not the problem, it's the lack of attribution. Unsurprisingly, these companies neglect standards of academic honor and attribution. Some human researchers also used to do that but in a discipline like mathematics this used to be a small problem because people tend to be so specialized that very few people could just grab someone's research and quickly piggyback on it, and if they do, colleagues will generally understand what happened. Unfortunately, AI is changing this.

You are both using a different definition of sharing I believe. When people have an expectation of privacy, use by others should be forbidden. Tech has gone completely off the rails with the use of private data.

He was arguing against two years. A lifetime is fine.

I have the setting turned on in Gemini Pro even though I work on proprietary code because 1. the setting allows for some (very limited) "memory", and 2. I consider my source code almost public even when it's not open source because I don't work on programs that involve extremely high level of know how or proprietary algorithms. It's mostly CRUD that can be copied in a myriad of ways, whether people use my methods or other methods.

If Gemini can improve based on my code and sessions (maybe doubtful but who knows) and others can benefit from it, that would be a welcome side-effect.


It seems completely trivial to feed sessions to their own LLM and ask it to look for various things in them, from detecting problematic use cases to finding interesting mathematical work.

let's say that they ask a single question for each session they get. they are immediately doubling the compute they need in processing and then post-processing the same session twice.

nothing trivial about it. not saying they cannot feed "their own LLM" saying it isn't trivial especially at scale.

if you do not trust me try it without the "at scale" part.


It's trivial and a solved issue for the companies developing frontier AI models. Obviously, you don't even need AI for searching every prompt every user has ever written to find interesting topics, but you can create automated summaries and use AI on them if you want. There is no "scaling issue" here for companies who are used to processing almost everything that has ever been written anyway.

I didn't want to insinuate that it's trivial for small companies or individuals to do big data mining at that scale, sorry if I made that impression.


I not only think it is not trivial, I know it is not solved.

I don't trust your judgment, it's in my opinion even hilarious given that we're talking about companies worth almost a trillion dollar (4 trillion in the case of Google). Be that as it may, it was nice chatting with you!

I also find it hilarious. I wonder what will happen if they don't solve this.

That sounds a lot to me like a lie.

Based on publicly available evidence, we should believe the former because AI models have made continuous advances that have been measured. The idea that the continuous advances will stop currently has no evidential support. I'm not saying it's not a possibility, just that it seems unsupported conjecture right now.

By the way, there seems to be a new form of AI skepticism emerging in the US that comes from general opposition to data centers, and in my experience the AI skeptic part of it is wholly irrational. I've met people online who suggested, without providing any evidence, that AI is useless and nobody wants it. That's a very implausible take.


Who doesn't fear that...

If I was the director of an agency of the size of the NSA and was evaluating the options purely from that perspective, I'd aim at creating a file on every living citizen on earth, including their social network topology and their activities. Basically a Google search engine that includes information not publicly accessible. I'd create much larger files for persons of interest and authorize targeted surveillance of them, of course, but with today's means to collect data a complete world database on every living and many dead persons is well within the technical capabilities. It also makes sense and is rational, if you put aside moral considerations.

That's how I evaluate these things. If it makes sense and can be useful, it's likely going to be done. Notice that there is no law against this in the US if you exclude US citizens. It's perfectly legal and within their mission parameters to do it for non-US citizens. I used to think my judgments were a bit too much on the paranoid side but when Snowden published his leaks it turned out that I was roughly right about every capability the NSA had except for their internal security.


Yeah, I'm sure some system like that exists, although I'd assume that would be more in the CIA's purview. I'd be surprised if they kept a broad swath of data for most people though as the tech companies already do it and it's constantly up to date. If needed, a fed lawyer can work through the FISA court and the tech companies are obliged to provide the records.

According to the information I have, the CIA is unlikely to be involved with SIGINT of that type. It's just not their role. I agree that most of the information the NSA might collect will come from publicly available sources like data brokers, particularly if US citizens are involved. However, what I was talking about concerns real-time capabilities and predictive power, it's very different from targeted surveillance and anything involving courts.

Isnt that basically Palantirs business model?

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: