aboutlogic #20 | Can AI Prove the Riemann Hypothesis? | Tudor Achim (Harmonic)
Show notes
You can now become a channel member on Youtube. https://www.youtube.com/@aboutlogic
Your support helps us keep these conversations going! If you’d like to contribute, you can buy us a coffee here: https://buymeacoffee.com/aboutlogic
Get the HoTT Book for free (no advertisement): https://homotopytypetheory.org/book/ Thorsten Altenkirch: http://www.cs.nott.ac.uk/~psztxa/ Deniz Sarikaya: https://www.denizsarikaya.de/ Creative Production: Jan-Niklas Meyer: http://www.jammos.com/
AI #Mathematics #Harmonic #Aristotle #Lean #ProofAssistants #FutureOfMath #RiemannHypothesis #aboutlogic
Show transcript
00:00:00: How could this change the way mathematics is done?
00:00:03: I mean, are mathematicians going to be redundant at some
00:00:06: point?".
00:00:07: You start right away saying we want to prove the Riemann hypothesis.
00:00:10: Uh...I
00:00:11: think a bit of luck will be helpful but generally speaking i think that the tools are there!
00:00:20: Great welcome everybody to this new episode of About Logic and today we have another very special guest.
00:00:26: so tutor We're very happy you join us And your colleagues as harmonic doing great stuff.
00:00:35: One unicorn start up getting quite huge, quite quickly.
00:00:39: Working on artificial as mathematical intelligence and your to improve ours.
00:00:46: total was able to do quite some remarkable things.
00:00:50: so we're happy that you are here.
00:00:52: And Thorsten You have the first question right?
00:00:55: Oh yes So maybe you can tell us a little bit about our is total before We go further
00:01:01: sure.
00:01:01: well It's great to meet you both, thanks for having me on.
00:01:05: Aerosol is a system to do formal reasoning in lean.
00:01:09: we apply it both to mathematics and code And We actually have been focusing on lean for a long time because Harmonic actually started with the thought experiment back in twenty-twenty three.
00:01:21: So we asked ourselves and back then You know AI could barely do high school math or let's say even middle school math much less But back.
00:01:29: Then we kind of looked at the trends of progress And we asked ourselves a question, which was in the future when you ask a model and ten years let's say to prove the Riemann hypothesis.
00:01:42: It would think for awhile on probably give you one hundred thousand pages of text?
00:01:47: I actually felt that you might as well throw out trash because first off all there is some error somewhere.
00:01:52: it will be very hard to find.
00:01:55: even if theres no error um...it'll be really hard.
00:01:57: people understand its.
00:01:59: like alot of texts can't make sense.
00:02:02: To me, it almost seemed like giving a hue and being you know the Google code base but no access to compiler or type hints.
00:02:09: our signatures are anything.
00:02:11: So that's a problem.
00:02:14: And then also if he asked again would give your different hundred thousand page proof because keep asking over an over?
00:02:20: We thought there will be deluge of AI slop.
00:02:24: essentially That wouldn't be very useful people.
00:02:27: so we thought If you could formally verify the proofs you'd have a system that could do very hard math, but also make it accessible to people.
00:02:37: Of course we would be correct because this is formally verified.
00:02:40: We can talk about like recent exploits in the kernels But generally speaking those will go away And It's a finite code surface to check.
00:02:48: Um...but Also we thought that might make reinforcement learning more efficient.
00:02:54: So back then everybody was doing reinforcement learning with human feedback.
00:02:57: A lot of it was optimized towards chatbots and we felt if put AI improvement systems in a loop with the formal verifier, you might be able to get a very advanced mathematical reasoning system without spending tens of billions of dollars on it.
00:03:12: So for those two reasons we went all-in on formal verification.
00:03:16: now Aristotle's available is again essentially large language model.
00:03:20: reason writes lean and also informal mathematical reasoning And end product is Lean code that's used by mathematicians computer scientists and some cases software engineers as well.
00:03:32: Yeah, so how could this change the way mathematics is done?
00:03:38: I mean our mathematicians going to be redundant at some point or should they just changed their work and any views on this.
00:03:48: Well i think it depends what people think of the role of Mathematics as well as the demand for it.
00:04:01: Figured it out and We're at the end of math, And we have all the useful math that's needed for understanding The world in the universe.
00:04:10: then I think That there are issues right because AI systems Are quite good at Math.
00:04:13: Mm-hmm i Think they're arguably better than some human mathematicians.
00:04:17: In Some ways i mean They're obviously getting Better mm-hmm.
00:04:21: But I think that what's gonna happen is actually closer to What's happening in software engineering which Is AI is actually driving higher demand for excellent software engineers because There are high returns to leverage To acknowledge and the ability to use these systems into direct them.
00:04:38: Do they ends that human organizations have?
00:04:41: So if instead we think of math as something where there's infinite demand for it, And then we always want to do more and more advanced math Whether its understand physics or create math That might be useful in some way In a future thats highly unpredictable now Then I actually think the demand for mathematicians will stay stronger increase.
00:04:59: The kind of math people are, we'll do is different.
00:05:02: There would be a lot more directing systems understanding how the math fits into the rest of humanity's collective consciousness but... ...I think that's likely to world were going towards and in that world i'm very skeptical that demand from mathematicians goes away.
00:05:18: But I did thing.
00:05:18: we're gonna change the kinda work We Do In the same way that software engineers are changing the Work They Do.
00:05:25: And what would you advise a PhD student in mathematics?
00:05:30: I mean, how will they need to change the way their work given that there are two like Anastasia and Lin.
00:05:39: Well, to be clear...I'm not a mathematician!
00:05:41: I did undergrad math.
00:05:43: Um, I would really hesitate to give advice to a math PhD student who knows a lot.
00:05:47: They've probably forgotten more math than I know.
00:05:49: um But what i will say actually and it's something related To what you know?
00:05:52: What we value internally at harmonic when it comes to AI because We obviously write a lot of software with AI is I think that It's important to be open-minded And curious about this new technology rather then being scared Of it.
00:06:05: yes, I'm reticent to try it.
00:06:07: so What I would encourage any phd student to do Is simply You know not assume anything about its level of capability or it's thread to the future profession and just try It.
00:06:18: give a qualifying exam questions.
00:06:20: Give it research question See what the strengths and weaknesses are?
00:06:23: And most importantly see how they can employ for their own academic goals.
00:06:28: I think once you really get your hands on it, you realize that first of all it's very Very impressive but second of all, it's it's not really magic.
00:06:34: It's just another tool in your toolbox and it'll have a big impact on the future job But It's a tool that everyone is going to be using.
00:06:42: So I think the most important thing, it has been open-minded and curious about rather than afraid of
00:06:46: it.".
00:07:00: formal theory improving in formal mathematics.
00:07:03: And it seems to be that this is a perfect marriage, right?
00:07:06: I mean as you said how could one trust the output of this Riemann hypothesis Prover if not... If we don't have something like a proof sys checker here so yeah but maybe It's just important for mathematics But also further may be economically more critical applications of AI.
00:07:33: Yeah, we also see a big verification gap in software.
00:07:38: so I think for a lot of software development you know like let's say your building a website or...I don't know like an API for hailing taxis or something right?
00:07:51: In a lot these domains if your code goes wrong it is actually okay to simply restart your service.
00:07:57: If your database goes down and you restart then catch up from the log that was written And you resume some sessions and the user has to wait for ten seconds because something went wrong, then they resume.
00:08:10: So in these areas I think that what we're seeing is that vibe coding has been unleashed.
00:08:16: if you can tolerate mistakes by just doing something simple integrated into user experience a little bit You will see AI writing alot of code instead people.
00:08:27: But i think more safety critical area where failure as consequence somebody might lose their life An airplane failure happens or lose a lot of money because the bank gets hacked.
00:08:37: I think they're, The standard for correctness is much higher.
00:08:40: and what tends to happen Is that?
00:08:42: People actually have to review the code very closely.
00:08:46: And if you are at all familiar with Amdahl's law from optimization computer science Typically when you optimize one part of a process You don't get through but improve in the expected Because there was another bottleneck which was big part Of the process before.
00:08:59: So i think that in safety critical coding Maybe the programming was half of work and verification is half.
00:09:06: So with VIBE coding, even if you cut half down to zero which is the programing part You still only get about a fifty percent speed up.
00:09:14: And so I think that If you have tool That can actually verify The output of VIBE Coding in other tools like that Um...you might be able To build more important software than before.
00:09:24: One part for sure Is formal verifications Essentially looking through every execution path Of a program.
00:09:31: Nothing bad happens, it matches the user intent.
00:09:34: And generally speaking you avoid the kind of software catastrophes that we're seeing pretty often these days with a new generation of AI attackers?
00:09:45: Maybe can I ask something like going one step back.
00:09:48: for the audience who haven't read your papers yet and don't know harmonic... These LLMs used to be very, very bad at math before Lean came in.
00:09:57: We all remember asking them addition questions but they got it wrong.
00:10:01: And now, I mean there were some successes from you.
00:10:04: From other companies and would your mind are giving us like a very close short timeline?
00:10:10: What was already formalized with ours total?
00:10:14: how did it succeed in the IMO international math Olympia so to speak?
00:10:18: yeah maybe what we're working on this course?
00:10:22: no We started twenty-twenty three.
00:10:25: back then models were horrible at math.
00:10:27: we worked really, really hard and in twenty-twenty five.
00:10:32: We shared gold with open eye and deep mind at the IMO.
00:10:36: And one thing that was different about our solutions is they didn't have to be human checked.
00:10:41: So because they compiled and lean we could be certain that The uh...the solutions are correct.
00:10:49: That didn't matter quite as much at IMO level problems Because it takes an expert human maybe thirty or forty minutes.
00:10:56: Check something in geometry actually can be a little worse because like the proofs are a little more intricate and annoying but generally speaking you can check them under an hour.
00:11:05: But as you see, they capabilities of the models improve.
00:11:08: You know They start to do a lot more advanced math.
00:11:10: So in November I actually think we were the first to solve air dish problems with AI.
00:11:16: And then hindsight, you know There's there's a wide distribution difficulties.
00:11:19: so The ones solved in November tended to be on the easier side.
00:11:23: pretty big wake up call because I think Aristotle was the only tool in the market where you could just give it a problem at that level and then we'd work on giving solution.
00:11:32: And crucially, It was verified.
00:11:34: so generally speaking know You can trust the output?
00:11:38: We actually saw and still see a lot of users who will take the output for another model.
00:11:44: run through Aristotle double check if reasoning is correct and often Aristotle finds subtle issues fix them even updated proof.
00:11:52: So, twenty-twenty three we started July.
00:11:56: twenty-five IMO gold.
00:11:57: November.
00:11:58: twenty five Airdish problems right pretty big in improvement.
00:12:03: and then since Then there's of course been a steady set of improvements from both harmonic as well other companies.
00:12:08: Um, i think most recently the most impressive remote I saw was Open AI solving those ten Problems at least I've seen publicly.
00:12:15: You know these are like pretty serious research questions.
00:12:19: they're very real.
00:12:20: um, I think They also formalize them in lean And I think that it's very clear.
00:12:24: we're on an exponential trend of math capability.
00:12:27: Although, one thing that i'm really proud and excited about is the fact that... ...I Think That We Basically Caused This Phase Transition Where Leaving Us Had The Difficulty Of Math That AIs Could Do.. ..I Think Aristotle Showed For The First Time It Is Really Possible To Formalize Serious Research Math.
00:12:49: kind of both referring correctness as well, is impact when it comes to papers.
00:12:54: So we hope for a future that's a lot closer to software open source where all the math has done in public perhaps on GitHub or other platforms.
00:13:04: everything is shared as lean proofs.
00:13:06: so you can just tell correctness immediately from The fact that the code builds but importantly could have new collaboration models.
00:13:12: were You could fork repositories?
00:13:19: And this is a much more decentralized architecture for math than what was there before.
00:13:24: I think that's a very interesting outcome of the whole set of events versus the obvious point, which is that AI gets smarter over time as they do in many areas.
00:13:34: Maybe one small detail just about Aristotle software?
00:13:38: In the beginning you also used computer algebra software but now don't anymore need that right if we call correctly...
00:13:49: Well, the Aristotle agent that runs it.
00:13:51: you can use any tool at once.
00:13:53: It tends to not really need tools.
00:13:55: At the IMO we had one specialized solver for geometry.
00:13:59: This is for playing geometry.
00:14:01: Okay
00:14:01: That was far?
00:14:03: It's just a detail of the Imo and its'nt part of the aerosol toolkit anymore.
00:14:09: But generally speaking You know...you just need very simple tools And smart reasoning To do well in math.
00:14:14: I mean...Hero Math Editions working on chalkboards.
00:14:17: So clearly don't need too many.
00:14:20: Actually, as somebody who comes from formal proving and developing type theoretic systems.
00:14:26: And so the progress long time was rather slow because these systems like Lean and others are like Rock and Akta.
00:14:36: they're really hard to use.
00:14:40: They were a domain of some expert users.
00:14:42: you needed someone that knows their way around And the AI tools make these formal systems accessible because the formalization effort is very much reduced.
00:14:59: There's another perspective?
00:15:01: Yeah, absolutely I think... The promise of formal methods has been widely recognized since the beginning of software engineering and computer science.
00:15:09: The dream has always been to be able to write code and then prove its correctness with some assumptions in models.
00:15:16: I think the fact that AI can now write proofs of scale is exciting and i think will make it much more accessible.
00:15:23: If you ask someone, hey would you like to assure your code over a hundred percent correctness?
00:15:27: Of course they'd be like yes um The problem was until five years ago...the price tag was million dollars in three-years.
00:15:36: Ten minutes!
00:15:37: And the dollar then?
00:15:39: uh It's gonna be a lot more widespread.
00:15:41: We really looked at toy problems proofs that prime numbers are checked correctly and stuff like this.
00:15:48: Because it was clear that writing the code is top of an iceberg guide, I mean writing a code is at least linked but then verifying or writing verified code yeah?
00:16:02: It's Muhannad Time, Sala right!
00:16:06: And only tools like U-Tool makes us accessible.
00:16:11: Yeah see you next
00:16:12: time.
00:16:12: Maybe one little aspect you already hinted at.
00:16:15: I mean, there was a lot of human leans already right?
00:16:19: These libraries people formalized stuff and some of them advanced but it didn't scale yet.
00:16:24: Of course they're totally right But these tools were developed for humans so to speak.
00:16:32: Harmonic also donated quite a lot to lean.
00:16:34: maybe push the software more because now come these LLMs which are somewhat powerful in some sense than humans brute forcing and trying a lot of things in finding bugs in the kernel.
00:16:47: You mentioned them, is there some thoughts where this will co-evolve?
00:16:53: Will we need to push formal tools further in order to be able to really check the LLM output?
00:17:01: I mean it's still much more sure than when i'd write a proof down its ninety nine point nine nine nine percent.
00:17:08: but you get these final
00:17:10: epsilon.
00:17:10: yeah I think probably within six months we'll have a verified lean kernel, so it itself will be verified for correctness and will be passed these mistakes.
00:17:22: We see...I Think that's really important to invest in tools like Lean.
00:17:28: And i think Lean is by far the best language for formal verification.
00:17:31: It combines A very nice kernel Nice type system with tactic metaprogramming That's very convenient and easy For humans and AIs To write.
00:17:41: If you leave aside just the lean language itself, though I think that there's a very wide scope for human mathematicians in projects like Mathlib.
00:17:49: Um...I think it behooves those project to maybe move away from caring so much about how the proofs themselves are written instead of making sure the definitions and structures are correct set up well organized and comprehensible.
00:18:08: Ultimately, what we're gonna want as humans is to be able to assure ourselves that we understand the results AIs are obtaining and then they're aligned with our interests in math.
00:18:17: So I think you need to elevate their reviewing... The human review of method from proofs-to higher level structures.
00:18:25: but there as well LLMs can help when you discuss your view.
00:18:31: this LLM it's off
00:18:33: makes an easier
00:18:34: faster more efficient.
00:18:36: Ultimately, I think that humans...I think there's a choice for people whether or not they want to comprehend.
00:18:45: approved fully because Lean does let you verify it.
00:18:48: So you know its correct.
00:18:49: You don't need to comprehend it understand.
00:18:52: but where i would draw the line is out say that Humans absolutely must completely comprehend The abstractions at the elements of using.
00:18:59: and so while ellens may help for That actually thinks that kind of burning your own thinking tokens To make sure definitions are set up correctly and reflect the will of mathematicians.
00:19:08: I think that's a role, it is probably better left to humans.
00:19:13: Yes however when you try to understand the proof, an LLM is not just something producing an output but your can interact with it.
00:19:24: You ask for explanation and you can interactively explore...I completely agree.
00:19:31: obviously We know it's correct, but also we know why its correct.
00:19:36: I mean that is an aspect of proof.
00:19:39: But i see the whole of LLM not just in saying yes It´s okay... ...but being able to interact with a human and provide more explanation.
00:19:52: That one thing im using LLMs for.
00:19:55: I read paper then idk what was this definition?
00:20:00: How does this work?
00:20:00: You know, and you can use an LLM then to explore it.
00:20:05: To understand the content
00:20:07: here.".
00:20:08: Yeah I think LLMs are great tools for essentially exploring verified code.
00:20:13: so if a verified code is math-proof they're really good at poking around in understanding.
00:20:18: yeah like he said how's his definition used?
00:20:20: what was this theorem saying that kind of thing.
00:20:23: I just think its important have formally verified artifact.
00:20:27: thats grounding whole discussion just to guarantee that you don't have the system hallucinates and explanations or definitions aren't there.
00:20:36: Because no matter how good the models of God, they still absolutely will hallucinate things like
00:20:40: that.".
00:20:41: Sure!
00:20:42: And that's exactly a point producing some formal output because it goes at end of hallucination right?
00:20:49: But here I think if quickly mentions this isn't the problem now – How much do we trust
00:20:56: Lean?!
00:20:59: important so much for mathematical applications, but for safety application.
00:21:06: I mean if there is any issue in the Lean Prover and AI which wants to exploit it will find that right?
00:21:16: And then tell you a wrong malicious code.
00:21:27: The lean kernel is maybe eight thousand lines of code.
00:21:30: If you really strip away the auxiliary ancillary functionality, I just think that there's a finite list of possible bugs in that code and as we put more models against it those bugs will go away.
00:21:44: but importantly one can write a lean kernel in Lean itself to prove its correct.
00:21:52: people are working on that like Mario Carnero And once those efforts converge, then I think that you're done.
00:21:59: So the kernel is correct and it will only accept valid proofs... ...and we don't have to worry about this stuff
00:22:04: anymore.".
00:22:05: Yeah?
00:22:07: Okay!
00:22:09: I'm interested in- I think the lean kernel isn't actually lean enough because it has still- What's
00:22:16: C++?!
00:22:20: It's too complex.
00:22:22: This was my own research.
00:22:25: much smaller systems, you can write on a big one page and not thirty thousand lines of code.
00:22:31: And you can reduce something like lean to this.
00:22:34: so I think that there is way to leverage even more.
00:22:38: have really lean... The idea for lean was to have small kernel but in my point of view lean doesn't get very small kernel still quite big.
00:22:50: There's
00:22:50: probably a bunch of ways to get at this, but I think what you're observing is that there are lots people interested in getting into a lean kernel where everybody thinks it's correct.
00:22:59: And whether its verifying the kernel itself or having simpler versions of it...I think we'll go one way another.
00:23:05: No no!
00:23:06: That isn't an alternative.
00:23:07: i agree Verifying the Kernel itself Is a way to go.
00:23:10: and just think
00:23:12: How?
00:23:12: To verify The kernel itself?
00:23:16: I mean, maybe just one side remark.
00:23:18: I really like how human it still looks like.
00:23:20: we have systems looking at other systems cross-referencing each other so It's not about one foundation where?
00:23:26: We just add all the trust but it's still like peer reviewing on On the formal head of the spectrum said to speak with different mechanisms increase a likelihood of success.
00:23:39: Maybe one slightly different area.
00:23:42: Normally in formal math, the goal shifting happens from the philosopher.
00:23:46: you say we have gold level IMO and the philosopher says yeah but that's not research.
00:23:52: yet they solve adage problems.
00:23:54: And the philosopher said this is special math or it looks like that We have annuals paper-level results.
00:24:01: They still aren't happy.
00:24:04: You start right away saying want to prove the Riemann hypothesis the biggest open question, so to speak.
00:24:11: Is there like one key ingredient that we need to develop before League really getting into this most ambitious goal?
00:24:18: There is concept development with AI or something like that?
00:24:23: Or do you think things are already there and it's just a matter of compute our time or
00:24:27: luck?".
00:24:29: I think a bit of luck will be helpful but generally speaking i think Continue to make them be able to reason over longer contexts and longer trains of thought.
00:24:43: I think it would be difficult to bet against the current paradigm Achieving that level performance.
00:24:49: We actually collaborated with the American Institute of math recently.
00:24:52: It's a creative benchmark instead of driven by AI companies one driven by people.
00:24:57: where we got together And we said well look here is set of problems that mathematicians agreed as a group.
00:25:03: You know if AI can do these then you know, it really means something for human math.
00:25:08: And I think that's kind of an interesting social point, which is that you should let communities of academics whatever field they're in get together and decide as a group what does AI mean for them?
00:25:20: That could be deciding how to use it.
00:25:23: How do accept it in publication?
00:25:26: but also importantly how assesses the group its impact.
00:25:32: If have process like that becomes accepted by you know, people that are like oh well it can only do the IMO or like oh we're going to air dish problems.
00:25:40: Or Oh he could only do double, you know a cycle cover conjecture.
00:25:44: um It's little harder too be credible when saying if entire field is agreed okay You know these certain steps where?
00:25:51: We think its impressive
00:25:53: One goal Kevin Bussard is formalizing The proof for March.
00:25:59: last problem What would u think about?
00:26:03: how soon will this happen.
00:26:04: I'm sure AI would help very much to speed up the project, but...
00:26:12: I think it should probably help us speed that up and I don't know details of what he's trying to cover But probably he sees at time lines are moving as well new eyes get released.
00:26:23: What about P not equal NP?
00:26:27: I think we probably want to find a proof first.
00:26:29: And, and i'm not sure if that might be one of the wierder ones...
00:26:31: But Riemann!
00:26:32: We don't have a proof either.
00:26:33: so hang on come out?
00:26:36: Yeah..I guess there may feel positive resolution for it but no- I think PversenP is an interesting one because like I am NOT a math addition my understanding was some of these Millennium Prize problems people do feel were making progress even though they are very far from the answer.
00:26:56: I think P versus MP, it's just really unclear if we're any closer to what then were a thirty years ago which is the tricky part of life.
00:27:04: There are no progress right?
00:27:07: So maybe... Well i'm
00:27:09: not an expert CS person
00:27:11: and one expert in P-not equal NP but an expert has told me recently that One issue with P-Not Equal NP Is as you say there is no progress.
00:27:22: yeah But may
00:27:23: be
00:27:24: AI will help us explore things very completely stuck.
00:27:29: I don't know...
00:27:32: Yeah, i think it would be interesting to see if AI can help there?
00:27:36: We'd probably get intermediate results before the full problem
00:27:39: though.
00:27:40: Sure
00:27:42: Maybe one question again with these human and AI interactions Are you aware of like one lemma or something that was created by an AI That's useful for AI tools?
00:27:54: maybe not that often used by human beings, like the analogous would be Sledgehammer in older theory improving software.
00:28:02: I mean it's smart pattern matches and is super powerful but it's not helping me as a human right because my brain is not wired like that?
00:28:10: Is there like vision of this divergence or...?
00:28:13: I think we'll have to see...I don't know if examples like us now but i also don't would comment on this, but I get the sense that there are papers now where a lot of the content does come from AI.
00:28:32: And in the same way that humans write theorems they're useful by other humans.
00:28:36: probably some of the papers coming out were the theorems or useful to other mathematicians.
00:28:41: those are created by AI and so that will be an example.
00:28:45: what you looking for?
00:28:47: Yeah actually i was using AI not as total Today because there was a problem.
00:28:54: I've been thinking for awhile and that thought it was too laborious to work out all the details And then i just used an AI To do it, because they couldn't be... I mean It could not have been bothered to do it right?
00:29:06: Well..it's just too overwhelming.
00:29:10: I think That is common use of AI even now Something very laborious.
00:29:17: You don´t want carryout all the detail But you give it to an AI.
00:29:22: We'll do it for you, what?
00:29:25: Yeah.
00:29:25: I hope that's the world we're moving towards.
00:29:27: In the AI world there is a lot of interest in replacing all human labor like can we replace cleaners or something?
00:29:38: and i just think thats kind of wrong way to look at it... Yes.. ...we can use AI to augment all human mathematicians ,i don't see why theres an obsession with this idea.
00:29:49: Yeah, I mean
00:29:51: in this case there was some work which i wouldn't have done otherwise.
00:29:55: Yes exactly
00:29:56: yeah and nobody was replaced.
00:29:58: it was just unfeasible right?
00:30:01: And then becomes comes feasible.
00:30:04: actually let me go to another topic uh which is close to my heart as how would tools like Aristotle and Lean be used in teaching mathematics?
00:30:13: if you got any any reviews on this or any thoughts on this maybe... You
00:30:20: know I'm sure that There's ways to plug it in.
00:30:22: I think some educators are trying, and there was someone at Berkeley that looked into this... My sense is as a civilization we've been educating kids for several hundred years if not longer In more of like a liberal classical sense And i think were probably not too far away from the optimal way To teach people stuff before they turn eighteen.
00:30:51: So it's really unclear to me if AI will on average make a huge impact.
00:30:56: On what students know by the time they graduate high school, I think however at the top end um It would probably be very big impact because right now in schools At least America There is not enough resources for kids that are strong at math or science and physics To explore their full potential.
00:31:15: And so from that perspective having super intelligent tutoring system is always correct, I think will have a very big impact on learning.
00:31:26: Actually let me say...I'm having problem because what i am teaching logic to computer science students and they are using tool like Lean or was using Rop before it's an abbreviation.
00:31:41: Because how do you check whether somebody understands?
00:31:46: Or prove something?
00:31:47: as you said writing not paper It's almost impossible to check whether students understand what they're doing.
00:31:55: But asking them, to prove it using software like Lean or any other systems makes a difference and make some interact... I mean in the end you really have to understand that.
00:32:07: but this is now problem because instead of really paying their head against the tool They just use an AI to do it for them, right?
00:32:21: So cheating is programmed in.
00:32:26: What's your thing about... I mean from me that's a personal problem.
00:32:33: Yeah!
00:32:33: It's tough.
00:32:34: maybe just live exams-I hear that a lot of universities are returning to oral exams in person
00:32:42: with four hundred students.
00:32:46: I think it's fair to ask, you know like maybe if students are cheating.
00:32:52: To that extent Maybe they shouldn't be in the class You know If they don't really want to learn at.
00:32:55: but i think for The core classes you kind of have to just assess people fairly By being in person and Just removing AI help?
00:33:04: I also Think That there is.
00:33:05: There Is a time And place right so you can Also make sure that Students Learn how to use AI by giving them exactly sizes that require AI.
00:33:12: But if we agree as University as a class is the society that we still care about humans understanding things then I think you gotta have tests to make sure they were also not regressing on them.
00:33:24: Yes, but actually all of when they learn it and in learning process i thing after appeal to them if ever use a forklift in the gym You may able lift all weights But don't train your muscles.
00:33:43: So that's the line here, right?
00:33:48: Yeah.
00:33:48: Although I think we have to give credit to students and they also prioritize their time.
00:33:51: so you know maybe they are really interested in some classes but less interest than others.
00:33:56: They're more prone using AI to complete requirements.
00:33:59: You notice something interesting which is that We interview a lot of people coming out from school And when AIs started working back early like late twenty three or so And what you would see is that the students kind of overused it, and they started doing pretty poorly in interviews.
00:34:23: But interestingly enough I felt like
00:34:25: self-corrected.".
00:34:26: So later on we noticed that students were recovering getting back to decent interview performance... You talk about some of them but then kids also adapt.
00:34:37: so for a while they thought okay let me just use AI for all my assignments.
00:34:44: they kind of realize that, stops them from learning.
00:34:46: And now I think your average undergraduate has a bit more nuanced view on how to use AI?
00:34:53: They do understand if you overuse AI will learn less and suffer some consequences down the road but it is giving opportunities for themselves.
00:35:03: which classes i really care about versus ones i care little?
00:35:06: And since you are on a tight schedule, is there any topic?
00:35:10: You would love to talk about something.
00:35:12: About the future.
00:35:12: we haven't asked you.
00:35:14: Well I'm curious why your guys' views of the future of math and CS in academia...I
00:35:21: mean i am philosopher by training.
00:35:23: for me it's always interesting part these back-and forth between the formal and human interpretation or their meaning shifts.
00:35:31: will things change There!
00:35:34: I think practice show AI tools will go into research practice.
00:35:43: It'll be a shift of focus towel, and you also said it today that this proof digestion concept development would become more important.
00:35:54: so I think the recent impact is what's interesting?
00:36:00: And it will shape math because it will bias us towards those fields where these provas work better or where there's more data.
00:36:09: But I mean, this is just descriptive right?
00:36:11: No normative stats on that.
00:36:14: a lot of things.
00:36:14: bias practice
00:36:16: Yeah
00:36:17: yeah and i've already sort given it away in bit.
00:36:21: So so I think That the greater opportunity now to build
00:36:27: formally
00:36:29: safe developed mathematics Which which was out-of-reach before And This I Think Is Incredibly Exciting And actually, on a number of levels.
00:36:41: So first off all, formalizing something is already quite hard.
00:36:47: I mean you have to use the syntax and so on but you can also use an AI To help with translation from English to the formal mathematics.
00:36:58: Now obviously now we must be able read it.
00:37:00: We need to reverse.
00:37:02: This is actually what I wanted to say, right?
00:37:06: You shouldn't rely on the AI.
00:37:08: But translation first of all gets a code... It's quite hard!
00:37:13: What are the syntaxes?
00:37:14: or how do i use this?
00:37:16: and so on?
00:37:17: but obviously there'a problem How you trust your own specification?
00:37:24: Or how do you trust any specifications?
00:37:26: because that's obviously the weak link.
00:37:30: Yeah,
00:37:31: I wonder if you know things will go kind of like sports where?
00:37:35: You have Mechanical tools like cars and boats that can effectively run and swim much farther and faster than humans But you know we still do it because we enjoy it.
00:37:46: If I think that math might go more in that direction although with much more applicability to the outside world um i Do think about.
00:37:53: You know if ai's really can prove anything better than a human.
00:37:59: There's the question What is the value of proof digestion versus understanding frameworks?
00:38:04: Because you know, if proofs come down to just techniques that you just repeat in certain cases and AI can digest it better than humans.
00:38:12: Maybe there's no point at a human understanding the proof but I do think so long as math explains world we care about explaining the world.
00:38:20: i think that humans will have to understand abstract framework.
00:38:24: If I
00:38:26: would go further, you need to understand even the little steps and rules of how things work.
00:38:36: Even once your understanding it can lead into an AI to do that for you.
00:38:40: but at some point in time you have to understand how logic works or precise reasoning works.
00:38:48: because i think you cannot learn how to judge proofs without being able produce them yourself.
00:38:55: I just wonder if in the future, it will still be useful for humans to judge proofs.
00:38:59: Like maybe we'll be a post-proof judgment world where It's only frameworks that matter and not so much details of implication chains from one theorem another.
00:39:10: The question is whether you lose bigger picture If never wrestled with muddy detail We will still
00:39:16: do that.
00:39:17: Oh yeah, no I think...
00:39:20: That's the question!
00:39:21: It is a bit like Harry Potter you know?
00:39:24: You learn very basic spells and then only when you master them are ready for more...
00:39:30: Great analogy Yeah
00:39:32: More like the Polly Juice potion or something you start with in Lingardium And then you put the Pollidoos' pulse on
00:39:42: it.
00:39:44: I think we can thank you for a lovely informative session.
00:39:46: If you don't have anything else, do really wanted to ask Thorsten or Tudor if he would tell it?
00:39:52: Anything important to say?
00:39:54: This is
00:39:55: great guys!
00:39:55: Thanks so much for having me on very interesting and fun discussion.
00:39:57: Thank you all for coming.
00:39:59: maybe i should say this is not sponsored And We don't got any question in advance.
00:40:03: So thank You for these honest interaction with us.
00:40:07: We are Very happy that you took the time and then hope to see you again at some point.
00:40:13: Yeah, maybe ours total gets the fields metal or something like that.
00:40:20: Guys this is great.
00:40:21: Thanks for having me on very interesting topics and thanks for making it late your time.
00:40:26: Thank you
00:40:27: sure of course.
00:40:28: great then.
00:40:29: bye-bye
00:40:30: Cheers guys thank you.
New comment