Seattle vs. the Bay Area, departments that ignored machine learning until it ate the world, and the long hunt for whatever algorithm the brain is actually running. This is the Master Algorithm conversation: Minsky’s “one damn thing after another” versus a few deep primitives still waiting to be found.

Where to build (Seattle vs. Bay Area); departments that ignored ML; n-grams to deep learning; the brain’s algorithm and the Master Algorithm quest.

Transcript

Pedro: I’m telling you, there’s stuff happening right now at, who knows, maybe at MIT, maybe at Pong University or some other. That is going to be a big thing. And again, going back, getting back to our earlier conversation, the Bay Area cluster is very big.

But it’s too homogeneous.

Pablos: I get it.

Pedro: You want to, you’re going to go to one of these because there’s dozens of them.

Pablos: There’s a, it feels like there’s a, maybe, I don’t know if you have a lens on this, but for AI research, a lot of American universities have been gutted, but Switzerland, meaning people are moving into tech companies.

Pedro: I wouldn’t say they’ve been gutted, but.

Pablos: All I don’t know why you, aren’t working at Google right now. It seems like they’re, trying to get everybody they can. But,.

Pedro: It looks to me like.

Pablos: There’s a hotspot in Switzerland and a hotspot in Oxford still. Do you, do you see any others like in AI.

Pedro: Research or. So exactly. There is this misperception. Sorry to call it that. I might have it. That industry has guided academia in AI.

Pablos: This is not true. Great.

Pedro: Because only a small fraction of the people go to industry.

On any given day. In fact, right now people have done these, measurements. If you, even now, If you look at the papers in the top conferences, There are more authors from academia than industry. The problem is that, for example, you see a paper that’s coauthored by three people from universities and one from deep mind. It’s like, Oh, deep mind has a new paper. You’ve got to, you’ve got to filter that out.

And, there’s a lot of people who have a leg in the university and in industry. People are like, XYZ has gone to go. They’re still professors.

And if they’re playing their cards this is actually good for everybody.

Everybody benefits. There’s a lot of people who are halfway between them.

The, and the perception that somehow now all hthe AI talent is in industry, there’s a, there’s more and more. And Google, at least they have certainly tried to quarter the market. They did at one point.

At this point, it’s impossible. You’re not wrong in that regard, but like the, if you look at the reservoir of AI talent in the world right now.

Academia is still the biggest one. And moreover, you could be a 17 year old kid to do, keep doing really good AI research and be discovered that way, even with no PhD. But again, going to grad school, finding your people, finding a good advisor to a principal, like why not do that? Go the hard road. I, to, there’s, a, at least a hundred, I’m making up this rough number. There’s this R1 universities, like, in America. There’s at least a hundred universities in North America. Canada is important. Actually. Great example.

Quick parenthesis. And I, I find this, the list, Where did the machine learning, let’s, look at deep learning.

If you had asked me this question, 15 years ago. And I said, like, NYU has a good group. The university of Montreal has a good group. It was like, what the heck is that? Montreal. Who’s in Montreal? There’s this guy called Joshua Benja. Toronto is a, is a top department, which people tend to ignore here, but it have a very good, Jeff Hinton was there. These places, I remember someone at Stanford saying, great guy, but this is outstanding research.

I won’t name him because there’s no need. But, he said like, oh, those places are either too far away or too far down. Meaning either in Canada or down in the rankings. My friend, where, did the deep learning and moreover, like when I would to go even further back and then, When I was, as I mentioned, like when I was looking for someone, so for someone to get a PhD in machine learning, the Stanford’s, the MIT’s, they had nobody.

It’s not that you only had, they had three people. They had nobody because they were like machine learning back then was their topic.

They were still doing. Another question is like, where, what are the, what are the, what are the topics right now?

That, that I could name a whole bunch of like weird topics in the places that are doing them. Another, but Schmidt Huber, of course, famously was in Switzerland and certainly deserves more credit than he got. And he’s very bitter about it. But like, back then there was like, he had the fortune again, the network effect comes in here. He was in Europe where neural networks were even more ignored than there were even fewer and far between people doing them. There’s Cambridge. There’s, there’s,.

Pablos: Wait, well, for you in the long, kind of in that long history of AI, what were the inflection points that gave you a sense of excitement or optimism? Did you feel it with, I probably, you got like GPT 2, I got access early on and I felt something there. Not necessarily unexpected, but I was stoked that it was finally working. What, do you remember any moments along the way where you felt like we, it’s really going to work or?

Pedro: Good question. There were a few. The first one in some ways, the most important was, this is 1986 maybe. And I go into a bookstore, this is in Portugal, And I see the bookstore almost by chance. They had the little, this was a tiny bookstore, the size of a living room in a little shopping mall.

They had a tiny section. I don’t know why of foreign language textbooks.

I know I have books that have to like electrical engineering and psychology books, but they did.

And I saw this one book there called artificial intelligence.

Pablos: All.

Pedro: It was the first ever AI textbook written by Eleni Rich published in 1983, just a few years prior, it was like a small format book, like a trade paperback, 400 pages. And I was like, what could this be? It seemed almost like an oxymoron, artificial intelligence. The second time I saw that, I used to swing by that bookstore on my way back from, from, college. And then I bought it.

And I read the book.

And, when I read the book, I felt at least, I felt two or three key things. One was like, wow, if you can do this, is going to be amazing. You could, I think I was going to say it would be hard to, but I would be wrong. Cause many people saw it. But me reading that book, I was like the potential of this is, extraordinary.

That was the first step. But also machine learning was, but at the same time, the state of the art was extraordinarily primitive.

And I’m like, there was some stuff that was The problem solving. They’re like, I see what they’re doing. This is my use, but this ain’t, what, these days back then the time didn’t exist, AGI. And that was always what I was driving for. That’s what I got interested in was, the potential of having not just, my goal has always been free intelligence. Part of what really drove me is like human intelligence is amazing. How it evolved is amazing, but it’s so damn limited.

Every day I’m, I come face to face with how stupid I am.

And if only I could be smarter. That has always been the goal. I saw that potential there. That was the first moment, but he didn’t, within that moment, there was another, which was like, they had a tiny little chapter on machine learning.

Clearly an unimportant topic.

That was an even more primitive, but I was like, this is it.

Without learning, AI will not succeed. We’re learning AI will go very far. Which at the time, nobody, even in AI believed in, except for a few people like, Jeff Hinton and the first moment. Right. Now, when I got into grad school, It was in the deepest AI winter.

I had colleagues leaving saying like, this isn’t going in. And then grad school saying. This is the mid eighties or something. A friend of mine saying like, machine learning. Like it’s impossible to do this now. But, I had to argue with him about this. And he left and I stayed. He said, machine learning is played out. We’ve discovered what there is to do. We’ve asymptoted. Wow. And then like, you have not. Here we are now. But like, so the second moment happened while I was in grad school was the beginning of the big data boom.

There was this. Like, it’s easy for me to go to like, see, I was about machine learning. I said one day it’ll change the world and now it’s happening. But there was something very important that happened, which I should have seen coming, but it didn’t, which was machine learning. And therefore AI is where it is today because of the flood of data from the internet, from the, in the beginning was the corporate databases, you can 1995,.

When the whole, it was like, wow, companies have all this data in their databases. And then, and then the internet boom began. And there was one thing after another that just brought us more and more data.

But when I saw that, I remember like forming you to, the next, my classmates who were graduating when, from typical PhD, there are no job offers in your field. You find a job doing something to being flown around by the, by the IBMs and the AT Ts and like, I was like, cause they want it like at this point. And I could see, wow. With, and I remember like, again, there was this conference that was just starting called KDD and knowledge discovering databases.

And I, and I, just saw like, happened to see like a proceedings or covertly papers. And I asked, one of my perspective is this committee member, chair of the department. What do you think of this? And he’s like, eh, this is weakly, scientifically poor, like something like that. Don’t bother with that. But I was like, and, I, she started publishing papers there. Cause I was like, I know this is in the primate state, but like the potential here, if we can take advantage of this data, this is going to be like fabulous.

And like, I know of course, like this is like catently obvious, but at the time this, was people’s, state of mind. And then, maybe to cut a long story, what I’ll give you an, yes, it’s true that like when you, when you, let, let’s do this in chronological order. In 2005, give or take, Jeff Hinton gave an invited talk at Ichikai, the big AI conference. I think he was in Edinburgh that year, fittingly for, Jeff.

Cause that’s where he got his PhD, if I recall correctly. And he gave a talk about this thing that him and a few people were starting to do on which he called representation learning.

And I was like, this is what we need.

There were a lot of hot things in AI at the time. And like, I wasn’t a big believer in, most of them, but I was like, this is what’s going to get us there is representation learning. And started telling people about it, like, like 10 years before the, so that was a moment of, now you see something that you hadn’t, there were things before representation learning that was interesting and so on and so forth. But like, and then they started to have these successes in speech, Which people remember, but their first success was like, I can see where this is going.

And I couldn’t see how we would get there, but like back in a time, like this is the two thousands when hardly anybody believed in neural networks still. I was like, I started getting into that.

I remember just to illustrate this point, having like coffee with Joshua Benjo at the ICML and I think it was Montreal where, he lives. And, we were joking about, this thing that nips should have become nips and then kips, that was the neural information, And then, 10 years later after the conference, we did a study, like we like to do machine learning, like what are the keywords most predictive of acceptance and rejection? And the keyword most predictive of rejection was neural.

Because if you were submitting a paper on neural networks to nips in 1990 something, you had no clue what was going on.

By then it was Bayesian information processing systems. And then it became kernel,.

These were some of the things that were hot. Joshua was like trying to get me into, so I’m interested in deep stuff. He said, like, I said, like, cause back then, I’m doing a little bit of it. You don’t have to sell me on that. I’m already sold. And he was like, Oh good. Like, how do you think this is all going to play out? And, I said to him the following, I alluded to this joke and said, and the next decade, meaning the 2010s is going to be dips.

Deep information processing systems. And, Joshua giggled.

Pablos: This was like, so out there. Good.

Pedro: Funny.

Not even, he was one of the few true believers. Not even he believed in this, but just to say like, at that point I had this thinking that this is going to be the next big thing.

And then of course, there’s the, the Alex, the Chachipiti moment, all those things are, important moments. But I would say the following to jump to the present. I couldn’t predict that something like Chachipiti would happen when it did.

But I would have predicted that something like that would happen.

And, for, a couple of reasons, but the main one is actually very simple is that language models, even pre neural networks, it’s this jaw dropping thing that even like, back for the first language model started in the fifties, People were trying to do machine translation and stuff like that. And, even the simplest ingram models.

A trigram model is already shockingly good at synthesizing text.

You increase the length of, my kid, Wrote a, Markov, there’s our so called Markov. He wrote like in Python, a simple little Markov model.

And mined a bunch of, journals with it.

And gave me a paper, an abstract entitled to read without saying what he’d done.

And he, and he asked me, it was funny because he liked some, I think gender studies, He’s a little mischievous. And I said, he said like, so what do you think of this? And so like, it’s the same old stuff. And then he said, my Markov, but he called it Markov, but my Markov, but generated it. And it was this. 2000 and this was of course, well preached at GPT.

Pablos: Impressive.

Pedro: Google’s big revelation was when they had a spell checker that was engram based.

And this just killed all the rules that Microsoft, that you probably know this story, why would they be able to do this? Because since they have much more data, they could go from trigrams to four grams to five grams. And what I’ve observed, Is that like, and this, I think very important for Microsoft with, a, with a big room and small data, it doesn’t look With a trigram, it kind of starts to look, but you can still tell. But at some point it’s generating whole coherent paragraphs.

And so immediately you tell yourself, if this thing ever get to, my, it’ll generate whole texts that are completely good. This is what’s happening now. The, barrier, Was that if you, require exponentially more data to increase the engrams.

And, but these people already actually most, again, people, neglected, but for me, 2000s, this was a key problem. Is it like, what you need is you need to be able to build models of language that have the power of an engram with large N while not having the data and computational blowup. Yes. And there’s a few ways to do this, but the key thing that Transformers did was exactly this.

A Transformers is really like an engram model that by gradient descent picks out which of those words to pick to and doesn’t have the exponential blowup.

Combine that with today’s internet scale of data and we are where we are. In a way that the chat GPT breakthrough was not surprising. Now, as I said, I could not have predicted that would happen then.

And it’s one thing to think that this is going to happen. And it’s another, and you can sort of like forecast some of the things and I did that in the last chapter of the mastermind. But like the whole way this is all playing out today is way beyond.

Pablos: I think it’s there for, at least for people who were deeply into what was possible with computers. And, I had some early experiences trying to work on AI systems, not research, but in trying to build things with, those, even just early expert systems and stuff in the late nineties. It was, it’s not hard for me to believe that this track, I think I always believed this track.

I just, I think the miracle of chat GPT really was just that they got a massive amount of compute. They had a conviction to get that much, that many resources aimed at that, at one algorithm basically.

Pedro: To be fair, there’s something to give open AI their due.

Google was already doing this.

The thing that open AI did, which, was, I followed this all at the time. I was like, this is definitely a great thing to do. I didn’t see what was followed. Was they started the post training.

Because a large language model by itself, Trying to predict the next word and therefore able to generate it. Will generate plausible texts. But again, that’s really where the harnesses began.

It’s like, it doesn’t quite do what you want. And there I,.

Pablos: Training is a harness in some sense.

Pedro: The thing is like, now I need, but again, context of the NLP research at the time, this was a very natural thing to do. Because it was like, they have all the different tests and you want to do well on all of them. And basically what they were doing, which was not, was just a natural step as, often is the case. It was like, we were pre trained again. Jeff Hinton was talking about pre training back in that talk in 2000 something for different reasons. This is often the case, but like you pre trained on this very large purpose.

You take advantage of it. And then you, can fine tune it to the different applications. And then you do very well on them using the small data that is available for them. And then what chat GPT, as you probably know, was just a way to gather data from people. That’s why it has a clunky name and everything. And they bet around how many people use it. And the highest bet was, I think maybe Greg Boxer, a hundred thousand. They never, but what they do is that at that point, they actually took, by again, piling on another, again, machine learning, not a handwritten harness.

It’s like, you’re just going to do this thing. I’ll give you two answers and let you rank, which is the best one.

Which is exactly what Google search engine is based on. This is already a technology that we know is fantastically powerful. They just combine two things that on their own are known to be fantastically powerful. The large language say, and this sort of like AB ranking type of learning. And then they kind of turn it into enforcement learning, which was a huge waste. And, but, that itself is ironic, but it’s a different story. But like, but adding that part is what really turned it into something that, people could find useful to do, all that they’re doing today.

Pablos: If you had to, do you feel a sense of, I sort of feel a sense of annoyance that deep learning has just sucked all the air out of the atmosphere. And that everybody’s just trying to milk deep learning for everything you can get out of it. And by extension, the other architectural things go with it. Whereas it feels like there’s probably real potential in radically different approaches to AI.

They’re, being neglected. Do you feel that way? Or do you feel like there’s, that we should be, I know you can see obvious cases like Jan Lacoon saying, I want to do world models because he’s seeing the limits of what we can do with LMS, that’s probably the poster child, but some version of that is probably true. And there’s even for language, we know this is a wildly inefficient way to do language.

We know your brain is, has some other better way we haven’t found yet. We know that it’s, it can do it, but it should be better. And we’re making it somewhat better, but should we be searching for the other ways of making.

Pedro: AI go? I would say I’m frustrated more than annoyed and honestly disappointed with my colleagues.

Particularly the academic ones, because our job is to not, the point of academia is to not fall into those things. In industry, I understand you have a shorter time horizon. They tell you, can do all this, Richard, but like everybody knows, that only lasts so long. You should be aware of that. But, so you’re exactly getting your PhDs and knowing more and more about less and less until we know everything about nothing.

AI, what we’re doing is more and more research about less and less until we’re going to be doing infinite research about nothing. There is, there is very much happening. And it really frustrates me that like all these resources, 99 of them are wasted. This is all stuff that, when the next generation happens, wherever it might come from, it’s all going to be a complete waste of time.

I understand industry doing it for the short term gain and But in academia, there is no excuse for doing that.

If you’re whining to me that you can’t compete with the NLP that they’re doing in industry, then you shouldn’t be doing that NLP. You should be doing some other NLP. Because as you say, AI is such a hard and vast problem. My biggest criticisms of like, the Dario Mothers of this world, like you have just no idea how much bigger than your brain AI is.

There’s this famous saying, which you can see in the new light today, which is that for most of AI’s history, we were the loonies. We were the lunatic fringe. As Jeff Hinton famously says, actually, let me put in my own words, the lunatics have not taken over the asylum. But we were like, these people are crazy. And one of my arguments was like, I forget, who’s way back, I think in the 60s.

If we were so simple that we could understand ourselves, we would be so stupid that we couldn’t. Right. Now, there’s a flaw in this argument, but there’s also a truth, which is, let me put it this way. There’s a famous school of thought in AI, like Marvin Minsky wrote a whole book about this, the society of mind. That is literally intelligence is one damn thing after another.

Stop trying to find the solution because there isn’t one. You just have to do one thing after another. By that standard, the current AI research is failing because they’re engineering, I actually do not believe in that. I do not believe in that because of machine learning. I have to, because back when I started getting interested in AI, like in my early 20s, I started making a list of the capabilities we needed to have.

And that list just kept growing forever. Which was, and the thing is, but with machine learning, you can actually have an underlying set of capabilities that then generates those in contact with data.

And that underlying set can be much, smaller. And we’re seeing that now. Look at all the things that Transformer can do. We need to figure that out. I would also say that the brain is probably more complex than it needs to be in many ways. But I really don’t think we have found.

Like if you look at neuroscience and computational science, which I, like to try to keep up with as much as possible. It’s impossible to not look at the brain and not think like there is, there’s something that this is implementing.

And it ain’t backprop.

What the heck is it? And there’s a good chance Jeff Hawkins famously wrote the best sell about this canal intelligence that actually inspired Andrew Ng to go into deep learning and whatnot, among other things. A little known fact, which is like, there’s this algorithm that the brain is running. We just need to rediscover it. This was also part of the motivation for my title, the master algorithm.

And whether AI is the Minsky version of one thing after another, or there’s a few really key things. The percentage of the community that is looking for this right now is keeps ever shrinking.

Pablos: Smaller.

Pedro: Right. Now to paint the positive note, Very positive note. This is a recent development in just the last really year. There have been all these new labs.

Jan is involved in actually more than a couple of them, but he has his, Ami one. And there’s a whole bunch of.

Pablos: Things,.

Pedro: The world models, Bob Lasso, and we could look at the pros and cons of each. But what is really positive for me is that there are a lot of people deliberately trying to do this. And the VC is actually at this point because they look at the history of open air and they say, like, if I bet in one of these labs, chances I won’t go anywhere, but it could really be the next big thing.

There’s actually people doing this, researchers, not enough in my view, and even funding for them. And again, unfortunately, the model that they have are like, I’m going to be like the pitch that these people are making to the VC. Some of them are my friends, this is not meant as a criticism, but it’s like they’re saying like, believe me, I’m a really good AI researcher here. My credentials, I have a track record and I’m going to discover the next big thing in AI. Not only do I have no business plan, I have no idea what the next big thing is going to be apart from who it is. Give me a billion dollars, please.

Pablos: Exactly. That’s what it seems like they’re doing. And it might even pay off. A billion dollars after you make the discovery.

Pedro: But here’s the thing, In the aggregate, that might be the case, but the new labs have a huge flaw is that they’re all in stealth mode. They don’t talk with each other. As a research system, it’s bigger. It’s better than a few big labs only, but it doesn’t beat the open academic system.

Links

Recorded on July 24, 2026