EDGE AI POD

The Skinny Transformer: Squeezing Gen AI into Tiny Devices

EDGE AI FOUNDATION

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 42:47

The future of artificial intelligence isn't just in massive cloud data centers—it's happening right now on the devices all around us. This insightful panel discussion brings together leading experts from major semiconductor companies and academia to explore how Generative AI is transforming edge computing.

What makes this conversation particularly valuable is the panelists' emphasis on practical reality versus future potential. While many assume GenAI requires enormous computing resources, the experts reveal that today's edge hardware—from smartphones to IoT devices—already supports numerous generative applications. The key isn't waiting for more powerful chips but rethinking how we approach model design, quantization, and specialization.

Danilo Pau from STMicroelectronics shares a fascinating vision of natural language interaction with everyday objects like thermostats, while Qualcomm's Evgeny Kuznetsov highlights how real-time translation and synthetic data generation deliver immediate productivity benefits. ARM's John Mark Yodis emphasizes that education and framework selection are more significant barriers than hardware limitations.

The technical discussion delves into cutting-edge compression techniques, with quantization advancing from Int8 to Int4, Int2, and even Int1 representations. Professor Huanrui Yang explains how foundation models can be specialized and pruned to maintain performance only in domains relevant to specific edge applications. This targeted approach enables capabilities previously thought impossible on resource-constrained devices.

Perhaps most exciting is the panel's exploration of unique edge advantages—proximity to data, sensor integration, and specialized hardware—that enable entirely new GenAI applications impossible in the cloud. Through orchestration across heterogeneous computing resources and domain-specific adaptation, the next wave of intelligent systems will distribute AI processing across the compute spectrum.

Whether you're a developer looking to deploy GenAI on current hardware, a researcher exploring new compression techniques, or a product manager planning your AI roadmap, this discussion provides crucial insights into what's possible today and where the technology is heading tomorrow. Don't wait for the future—generative AI at the edge is already here.

Send us Fan Mail

Support the show

Learn more about the EDGE AI FOUNDATION - edgeaifoundation.org

Introduction to Gen AI Panel

Speaker 1

This is the panel session for Generative AI at the Edge and first . So this is intended to provide an outlook on use cases specifically with Generative AI and also about the technology behind Edge AI and also we make the attempt to provide some outlook what's going to happen in the next three years or so . But let me first introduce myself and the panelists . So my name is Wolfgang Furtner . I'm a distinguished system architect at Infineon , based in Munich , germany , and I work on AI for sensors and actuators , actuators being mostly power , so also motor drives , and so we're also looking at generative AI . But we've been able to assemble a panel of esteemed speakers and I would like to introduce them one by one , just by name .

Speaker 1

So we have John Marco Yodis Yodis here from from arm first seat , and we have Evgeny Kuzov , who is so I could tell Marco is is a Jenny I engineering lead at arm . Marco is a Gen AI engineering lead at ARM . Then we have Evgeny , whom we all know , qualcomm and chairman of the AGI board , and then we also have Juan Rui Young , who is a professor at the University of Arizona , and then , last but not least , we have Danilo . Danilo Pau from STMicroelectronics , being a fellow at ST . So I would like to ask first each of the panelists to briefly introduce their touch points and affiliation with generative AI at the edge . So maybe in the order of sitting . So maybe , john Mark , you want to make a start .

Speaker 2

No , we can sit .

Speaker 1

I will also sit down . Okay , good , so maybe we should we do have five microphones .

Panelists' Expertise and Affiliations

Speaker 3

Should work . Yep , okay . So yeah , my name is John Mark Yodish and my affiliation with GenAI is related to the work actually I'm doing in ARM , related to the deployment of generative AI use cases across all the ARM processors , from cloud to the very edge . When I say very edge , I also about the software , the software enablement and the software optimizations and all the interesting quantization techniques that are necessary for this type of workloads .

Speaker 1

Okay , thank you . So then , Evgeny .

Speaker 4

Good afternoon everyone , and thank you for joining this panel . Good afternoon everyone , and thank you for joining this panel . In addition to my privilege to serve this community , I also have a day job , which is I work for Qualcomm and in Qualcomm I lead our AI center of excellence in the IT auto and also cloud . Very excited about this place , this space in general , agi , but now with GenAI coming into the play here , I am triple excited about this opportunity because I think we're getting into a perfect storm . The hardware tools are becoming mature , the models are becoming more and more capable and we are getting into the space where these two forces collide and we have a lot of exciting applications ahead that we're going to discuss at this panel .

Speaker 1

Thank you , Huanrui , please . Hello everyone .

Speaker 2

I'm Huanrui Yang and very honored to be here , so I'm an assistant professor , so my main work is doing research and actually I have started to work on AI efficiency before this wave of AI and now definitely it becomes more important as models are getting ever larger . So the research in our lab is mainly focusing on designing ways to utilize sparsity , equalization , decomposition of those model compression methods to make AI models run faster . And we also do something on the software hardware code design side , where our Mean interests lies , on reconfigurable hardware like FPGAs and also computing memory hardwares . Thank you .

Speaker 1

Thank you and Danilo , please .

Speaker 5

First of all , thanks , wolfgang , for the nice introduction on top of what you say , that my role in ST for the nice introduction On top of what you said . My role in STMicroelectronics , in system research , is about tools and algorithms for AI at the edge and I developed my career in STMicroelectronics , so always in system research , and that helped me to develop an attitude which I think was useful to serve even more productively the foundation in different roles . For example , I contributed on device learning , on neural architecture search working groups and the last one is the generative edge AI working group . And I have to say that the foundation , thanks to the community at large , really helped and gave fresh energy to invest more in this field which is very active and evolving so fast , faster than we can imagine , sometimes

Applications for Gen AI at the Edge

Speaker 5

too fast , but that's fantastic and give me especially the opportunity to learn more and more and also to see the future coming , which is quite important .

Speaker 1

Okay , thank you . So let me first introduce the rules for this panel discussion . So we want you to engage with the panelists from the start , basically . So I will throw in questions and I will start with one panelist to just set the topic and start the discussion , and then the other panelists will kick in , but also you in the audience , you're welcome and encouraged to also directly ask questions to that particular topic , and I hope we still make it in time . So let's start with the first question that I'd like to address right away to Danilo again . St is having a lot of sensors , compute , connectivity solutions , so what do you see for generative AI specifically ? What new applications , new opportunities are arising through this new technology ?

Speaker 5

Thanks , wolfgang . I think this is a very important question . First of all , a little bit of background . So I joined the tiny ML organization as a partner as a partner since 2018 and it was all everything about a fixed AI . So you can imagine , as a semiconductor company , time zero . You cannot have optimized products for AI .

Speaker 5

At the same time , ai comes to the edge and therefore the first implementation were about microcontrollers , standard microcontrollers , and it was very exciting because tiny microcontrollers , small memory , small computation capability , tiny , under 1 milliwatt , power consumption , neural network the two things really matched very well and and it was a huge stimulus on on on research and on solutions to be exploited in the products . And then the hardware acceleration came , obviously because we need for image , for example , image processing , and on the other side , as Wolfgang said , we produce sensors . So the beauty of a sensor is that it's a two millimeter side package and if you put everything inside , such as AI , it's a huge challenge in terms of power consumption I mean under microwatts , not under milliwatts , and the silicon area is limited and doing also on-device learning . There it's huge , it's a big challenge and nowand . What I wanted to say is that , with the tiny machine learning , essentially us and the community , and us in the community were able to help in moving AI from the cloud to the edge . That was fantastic , because AI was meant to be on the cloud . I mean on 2017 , google News knew that to support three minutes of automated speech recognition , the Google Cloud was not enough , and the 40 pews and such devices were developed and and today we succeeded in moving a I at the edge , but then generative AI came in , so we are back on the cloud .

Speaker 5

Definitely , there is no doubt about that , and and and that's the game for us as a community , as part of a GI foundation . Do we have solutions now ? No , certainly not . Is too soon , doing our vertex if you are lucky , two years , maybe three , and it's a lot of investment . So , but that doesn't mean that we should not invest our best energy to make happen generative AI at the edge .

Speaker 5

The real question I am concerned about is for doing what ? And certainly one answer is to narrow the scope of the use cases such that the model can be not under one milliwatt but can be maybe ter one milliwatt but can be maybe terawatt per watt very efficient in term of energy , such that we use processors , multi-core processors , hopefully with NPU new wave of NPUs so efficiently that we can map what I mean image generation , image processing . I did a survey 80% of the algorithms are focused on image , type of processing , various forms . Then we have short language models certainly 10% , 15% and then we have text to speech , speech to text and other things . And to do what ? I invested also in a couple of studies , which are published , by the way ,

Technical Challenges and Solutions

Speaker 5

and one which is quite intriguing is the following one I mean there should be a thermostat somewhere here , right , to control the temperature .

Speaker 5

So imagine that I want to speak with the thermostat like I'm speaking with you and I'm not setting a specific objective . I want to be comfortable in this room and I want the temperature to be I mean , I mean to be such that the gradient versus the external temperature is not huge , maybe six degrees of differences to be comfortable . But I don't want to tell all this detail to my thermostat . So actually , the thermostat should turn on AI agent , which capability to reason , to get the information it needs to set the objective . I said so to do what ? What does it mean To have a natural interaction with a device ? So if we are able to develop such use cases and to develop very energy-efficient hardware . Then I see a brilliant future for generative AI at the edge . Really brilliant versus years of development of fixed AI .

Speaker 1

Are there any other concrete applications in the panelist's view ?

Speaker 4

Yeah , I think , if I may add , from the Qualcomm perspective , we see Gen AI is a huge opportunity for us and for our customers and this is driven by two again driving forces the hardware is becoming very , very powerful , models are becoming more efficient and in a simplified way . I think about Gen AI as a productivity tool in general . And if you think about Gen AI as a productivity tool , it offers enormous opportunity . It's a productivity tool for tools for manufacturing . It's productivity tool for people to be more productive .

Speaker 4

Some simple examples , like I saw downstairs doing like translations from one language to others . That's what we also are hearing from our customers , because the workforce is becoming very multinational and you need to translate from one language to the other . You do it in real time . It's a simple use case , but it's a very valuable use case . And also we should think about productivity tool GenAI , not just for our customers but for ourselves . Like some examples using GenAI for synthetic data generation . It simplifies development of uh agi models big time . Using gen ai for coding , using gen ai for verification type of things . For for silicon design it's it's a productivity tool in a very broad sense . It actually you can apply it to every area in your workflow and get some benefits .

Speaker 4

And it's happening now .

Speaker 1

Okay , thank you . Any other thing to add , any other application that you have on your mind ?

Speaker 3

Well , I had a workshop this morning , right , and I talked about use cases and generally it's not something that will happen by something that can happen right now , and there already some interesting use cases around , for instance , the audio generation . I was mentioning about the audio jam for a DJ console . The DJ console , for me , is an edge device , by the way , okay , and when we talk about edge devices , we also need to remember that an edge device is at the edge of the network , so it's where you are very close to the data . It's not just about microcontrollers , it's also about all the embedded devices . The mobile phone , the laptop as well . The DTP , the tablets are all considered edge devices . The capabilities are already there .

Speaker 3

In my opinion , what is missing is actually the know-how or the knowledge to deploy these models , and this probably I can , because Danilo was mentioning about the compute capabilities . In my opinion , the compute capabilities is not a challenge . It's actually an opportunity , because we have already the opportunity to deploy these models . But , as we have seen with TinyML , at the very beginning , we need to educate , and here is the place where we need to invest . In this forum , we derive presentation to help the next developers to start on today's platforms and not on the next gen platforms , because there are already use cases that can be deployed . I add just one extra thing Gen AI has , of course , constraints on the execution time , but compared to classic ML , where if you don't meet a specific requirement you're going to fail , like object detection , here what you're going to affect is the user experience . What does it mean ? It means that you can deploy it , but even if it's not going to be super fast , the penalty will be mainly on the user experience , not on the fact that you cannot deploy it . Okay , good .

Speaker 1

So we're halfway already into the technical discussion , technical challenges that Gen AI is posing to us . So anything in that domain I mean ARM has contributed quite a lot to the ecosystem of Edge AI in general . But is there anything ? You say it's there , but is there anything still on the way to generative AI in terms of technical challenges , like on the memory side , on the compute side or so ?

Speaker 3

Well , certainly , the more compute , the better this will benefit the future use cases , but certainly we need to invest more on the open source projects , and we have already quite interesting projects . I would say many , many frameworks I mentioned this morning Executorge , lightrt , whispercpp , lamacpp , onanextra , on Time , and many , many , many more . So , of course , these are crucial because , although it looks like we have fragmentation , it's actually not the case , because each framework has its own strength right . What do we need to do ? We need to educate , we need to tell people where the strengths are in these frameworks and navigate the selection of the framework , because the framework selection will be part of the generative ai story any other opinions on blocks , roadblocks for jnai or technically challenges ?

Speaker 1

yeah , martin , thankfully , was volunteering to operate . We need another . We need one of the mics actually . No , maybe it's better to have the mic .

Compression vs Compute Capabilities

Speaker 6

Thank you . So you said the capability is already there and you were talking about edge devices are not just microcontrollers . Rightfully , these are laptops and all the other things . But then this opens the door to the question like where can you deploy the generative AI on the edge ? And this actually changes the paradigm completely , because the use cases you can handle for the microcontrollers or microprocessors , compared to the edge devices , if you consider even small GPUs as an edge device , you are not targeting the same application . And this way , if you start big , we have to start , but you are in a way , excluding a big community which is talking about the milliwatt or microwatt or these type of deployments .

Speaker 1

Yeah , battery-operated stuff , for example . Yeah .

Speaker 3

Well , certainly the I need a drink ?

Speaker 1

No , no , take your time .

Speaker 3

I can put you something while you talk while you start talking . If I start talking Sorry , so in the same time I can add for example if I start talking Sorry , you can stop , so in the same time .

Speaker 6

I can add , for example , something which can be very interesting is like having the possibility to do the generative AI for generating the anomaly data . The speaker just like who spoke here , he was talking about classification of anomalies . So we know like it's difficult to generate the anomaly data because you will not on purpose break your system to generate the anomaly data . So for detection I think it's easier , but generating the data on the device itself would make much better than anomaly classification in my mind , but at the same time you are saying it's probably too much . The second thing is for the models at the moment generate . When you think about generator models , we're thinking about transformers and all this . So probably there is some work needed on Having the topologies which are not transformer based or the topologies which are already Supported by all the big companies that are talking here about .

Speaker 4

There are many , many points you brought up . Maybe I'll try to answer if I remember .

Speaker 4

So some of them are on the heterogeneous compute you mentioned Like . There are some devices , tiny devices , then you have edge devices . They have a variety of edge devices , from phones to cars . They have gateways . How do these all things work together ? And I think that's how the future systems are going to be designed . It's going to be done all in orchestration of the things . Certain tasks , like the wake trigger detection , will be done at the microcontroller level . It triggers next level system and next level system and Gen AI will be also distributed in the same way and the orchestration becomes a big piece there , and connectivity to the cloud too . That's kind of what we in Qualcomm call hybrid AI . So how do you kind of orchestrate it seamlessly , all these different layers of compute , that will give you the optimal performance there ?

Speaker 3

So I like your consideration and , in my opinion , the two main ingredients for that are specialization and software programmability . Specialization , because this is the only way where , at the moment at least , with the existing topologies , we can tailor some of these models to perform a specific task , because these models , when we say these , are big models right , that's true , but it's also true that they're very , very capable to perform different tasks , to perform different tasks . However , if you specialize these models , you can do it for a specific task and maybe combine them together , like Afghani was saying , orchestrating these models to perform something bigger . The other important role will be played , in my opinion , by the software , because quantization techniques , for instance , will be played , in my opinion , by the software , because quantization techniques , for instance , will be super crucial .

Speaker 3

We are seeing not just Int8 . I mean this community , so we tiny amount the important role of Int8 . Great From this community . We had the Int4 , way before generative AI . Now we are seeing integer two , promising by a mixed data type fashion and including also integer one that could play an interesting role , although maybe the hardware is not there yet . What you have is a software programmability that can help to bridge the gap .

Speaker 1

Sorry , that can help to bridge the gap . Ok , thank you . Just a very small point , if I may . That probably is . Yeah , sorry , maybe you can also continue afterwards . We also want to give others a chance . Just one point , and then I promise I stop .

Speaker 6

So maybe , as you were saying that , having the possibility to do the shared computing between the devices . So we are not talking about cloud computing , but we're talking more like , then , edge AI plus more or less fog AI , where you have multiple devices and you need to do shared computing on the devices .

Speaker 4

Yeah , I think absolutely , and it was discussed at the panel discussion yesterday . How do you orchestrate these different devices ? Networks are going to be important , some advanced techniques , federated learning how do you kind of do learning in these different nodes ? And it's all driven by the use case . For some use cases , just like one device , you engineer it . For some of them , the whole factory needs to be kind of worked in this way . So it becomes an interesting opportunity for system integrators to connect all these pieces together .

Speaker 1

Now it's on you .

Speaker 7

It's kind of a follow-up , but maybe we can extend it . You mentioned quantization and there are other methods to compress and squeeze large modeling to small devices , but eventually it degrades the output . Squeeze large models into small devices , but eventually it degrades the output and sometimes also performance . You degrade the precision , which leads to performance degradation , and there are limits in some domains , as audio , as you already discussed on the workshop . Should we just wait and I'm shooting my own legs with that should we just wait that the compute capabilities will be able to contain Gen AI until we try to squeeze it more and gain minor improvement in very large investments ? Or should we just wait until the compute capabilities on edge devices would be better ?

Speaker 2

Maybe I can start on this . Yeah , so definitely , when doing model compression people see the degradation of performance . But people also see that if you start from a larger model typically like you start from a larger model , typically like you start from a larger model and compress to a specific size , then the larger the better point you start with , the better performance you will get after compression . So in this sense , actually a model compression technique has become more useful nowadays with all those abundance of large foundational models . And another aspect I think , like what they just mentioned , it's very interesting , is about the specialization .

Speaker 2

So one interesting thing that differs this current wave of Gen AI , this previous AI method , is that previously people just use one model to do one task . Then you compress it , you lose task performance . Now in Gen AI it's a foundation model that supposedly can do everything and when you compress it it . Now in JAI it's a foundation model that supposedly can do everything and when you compress it it may lose performance somewhere but still can retain performance elsewhere . And this gives you a knob in some sense that you can tweak it , that , for example , you just want , say , a robot to grab this cup and give you the water . Then you can just make the model to retain its performance on this specific task while potentially just lose some other totally irrelevant capability as you compress it . So these are the ways , potentially , that you can tune the way you compress it so that it still do your job perfectly but has to end up with a much smaller model size .

Speaker 4

YURI LITVINOVICHIY and to this point , just to be specific , there are techniques , as you know , that are fine-tuning , for example . They're becoming more popular and , from our experience , you can get a fine-tuned 8 billion model that behaves better than 70 billion model in terms of accuracy and performance for a specific , and that that because lms are general purpose hammers and you cannot use them for everything . It's better to want to use specific models , so fine-tune models for specific tasks other than lms , which are , by definition resources and then along for the best .

Speaker 7

I mean , there are some other than LLMs which initially started with overkill sizes , like more specific models which end for specific tasks . Should we further try to optimize them , accounting for the time it will take us the resource we will spend and the minor improvement that we might get ? Will take us the resources we will spend and the minor improvement that we might get , or just wait and say this current application doesn't fit this specific edge , and I agree with you that there are many edges , like devices .

Speaker 4

VYACHESLAV TYKHONOVSKYI yeah , I think that's a product question . It's always , I mean , in the product space . We always debate what is good enough , because you don't want to over engineer or under engineer . Optimization is always good , but it takes time , resources and other things

Industry-Academia Collaboration

Speaker 4

, so it is kind of a very tricky balance to do it um , yeah if I can

Speaker 5

because your commentary gets some some consideration from my side . First of all , maybe not all , but there are some industries I would say small , medium enterprises , but also bigger size industry that need to design their own models for a number of reasons , and if we go on the shop of the foundational models , maybe these are not what we need , what these people need , and therefore these people are tempted to design their own VQA , for example , and that starts to be challenging , very challenging , and then you apply optimizations , techniques which are known pruning , knowledge , distillation , all these sort of things quantization , post-training quantization all these things that we all know , because on fixed AI , that's unknown , and then you get 40% of accuracy , which means , first of all , it's already a challenge to spend time on these matters when you know that you can do developments on microcontrollers , with standard classifier , with standard detectors , and you are losing focus on these things and then maybe disappointed because you reach 40% . So , in conclusion , what I'm saying is that foundational models are not the way to go , maybe are not the way to start , are not the way to differentiate , are just models with weights but data sets . Training techniques are a black hole and this discourage a lot people on invest until this will be sorted .

Speaker 5

So certainly one contribution would be a tremendous efforts in teaching people how to design a VQA which is 40 per 80 percent accurate , not running on a microcontroller , because microcontroller , I'm sorry , these are not for generative AI at all , that's the bad news but for new processors which sometime are available Raspberry , for example , the PI and others . Maybe the performances are not super good , but you know , when industry understand how to manage application algorithms , then maybe they could design products . It was like that on fixed AI before to design and pews with without memory and so on . People wanted to understand , because making chips cost and doing for nothing is good for science but not for business . So I mean that's the point we need people to understand how to design their own workload , to train it , and foundational model I'm sorry , are not enough even for student-teacher models , and we need such an effort .

Speaker 1

Let's get another few points from John Marco , because I'm a little bit rushing , because there's also one other question that is important to me that I want to address .

Speaker 3

Yes , I'd like to . So you asked the question should we wait for the next gen processor or should we invest now ? Right , I think we should do it now because the capabilities are already there and sometimes we just need to see the things from a different angle . But let me share this story about the audio generation , stable audio , open model . Right , so the original model was trained to generate quite a long audio samples . Okay , I think it's over a minute , if I recall correctly .

Speaker 3

What was the approach to generate these audio samples on device in a reasonable amount of time ? Simply making shorter the audio sample . We don't need three minutes , we just need , for music producers or other people interested , a shorter audio clip , maybe 11 seconds , maybe 10 seconds . By simply adjusting that we wrote also on the paper we managed to make the model smaller . We could not use extreme quantization technique because , as you know , the floating point is crucial , especially for very good quality of audio , but actually was the key to unlock it . So what I say is I would like to maybe to explore now what is missing is the , an understanding for people to understand . What kind of performance can I achieve today on the existing platforms ? I think we require a bit of education on that in order to answer this question okay , how could I say no ?

Speaker 9

I'm just thinking , probably the reason why we're not able to deploy at the edge I mean generative AI at the edge at the moment , because we are thinking to borrow the traditional generative AI and put it in the edge , and probably we need a different way of thinking because edge devices are different from the traditional computers and like open AI . You know , chagbt came for a reason and they started the different way of thinking and then they , you know , took different direction from traditional language models and so on . So I'm thinking edge devices , for example , they have sensors , example , they have sensors , they have different kind of capabilities , they have mobility . We could do something like domain adaptation so we don't need to bring the whole knowledge of the world to the device , but we could do something specific which makes the model smaller .

Speaker 9

so I agree a little bit that current microcontrollers might not be able to you , of course , do the big models , but they might be part of the model .

Speaker 2

And I talked earlier today about decentralized models and so on .

Speaker 9

So if we have a different way of thinking , probably we need to detach from traditional generative AI and think of GenAI on the edge from a different point of view .

Speaker 4

Just a correction here that we are deploying gen ai image devices today . So it's happening , uh , and on the traditional non-traditional , I think there is a very interesting dynamics happens today if you think about traditional ai , like companies like meta working on llama type of models or hugging face working on small type of models , so they're basically using the same approach and apply them across the board . Ibm doing the same thing , they have tiny granite model , they have small models for edge model , specifically for edge , and then they have big for the cloud . But it's kind of the same family of model , same approach ,

Safety and Responsible AI Deployment

Speaker 4

using for design for different type of compute capabilities . So it's which is actually a very positive thing for this community . Thank you .

Speaker 1

Now , since we've warmed up and I think we could go on for quite some more time I still want to take the opportunity that we both have academia and industry on this panel . So I'd like to throw in that question of how can we better collaborate , or what would be the expectation for collaborating on this topic between academia and industry . And I would like to start with you , Huan-Rui , to maybe state what you do already , what your wishes would be towards industry .

Speaker 2

Huan-Rui Chen . Thanks for the question . Definitely now I think it's a very good time to collaborate between academia and industry Because , like previously , if you consider the time before , those large models actually , like people , are just doing a specific task and at that time academia people may just like have very clever idea to solve specific questions and the industry people take that paper and like visit through the engineering work to deploy it . But nowadays we find that there are two things that are very important for the academia that industry can provide One is on the resource and the other is on the applications or on the scenarios . So , resource wise , like as we are seeing larger and larger models , definitely we may need to have more data , more computations and so on and so forth to really scale up what we are doing .

Speaker 2

So , from the academia side , what we are doing it's more like general methods supposedly that can be applied to all different sides of models , but then we need to verify it whether we can , really can , and then the industry can champion there . And then , on the scenario side , ai are being applied across all different scenarios now . So there are a lot of applications that academia people didn't think about , but then for some companies they have , for example in Paris , talk about those monitors , like we can never think about that from our side . And then I think one good way is that we academia people knows kind of what can be done and industry people knows what needs to be done , and we can find the intersections there to build up this collaboration and see what new research area can emerge from those kind of interactions .

Speaker 4

I'll try to be short because I see that Martin is about to unplug us . I think that GenAI at the age is an amazing news , both for the academia and for the industry . Like , maybe two years ago , I heard a lot of pessimism from the academia , like we cannot contribute much to GenAI because A we don't have data , we don't have compute , we don't have money to do it , it's only the big companies , so it's the Meta , google . These companies can do leading-edge research there , they can hire us , but from the academia perspective we're just somewhere on the sideline .

Speaker 4

So now the center of gravity is moving from training to inference , and that that's an amazing news because you can get devices that you see downstairs for 999 . On that you see downstairs for $9.99 , $29.99 , and you can explore all kind of possibilities in terms of inference , what you can do with this type of device . It's again great news for the academia and that's what we see from the industry . It's also very positive things because , to answer previous questions about applications , if these devices are user-friendly , very easy for developers , then people are going to come up with all kind of interesting use cases for this type of device , like the hackathon , what Pete announced today . That's another thing . So it's again a great time to be in this , both in the academia and the industry , and the AGI is going to be a huge LLM , a huge driving force here which is driving the force here ?

Speaker 8

I was wondering if there's any way to prevent the generation of inappropriate or harmful content when Gen AI is run locally . Like Von Braun said that , the rocket can bring you to the moon or bomb the city , can we prevent the bombing of the city right in the metaphor way ?

Speaker 2

I guess this it doesn't matter if you are running it locally or online , like this is always an issue for regenerative models . Um , definitely , from the model designs , there are a lot of ways . People are trying , like I I can't say it's already being resolved , but people are trying I can't say it's already been resolved but people are trying to identify and stop the model from giving the appropriate responses . And for compressed models specifically , the model capability may be limited if you make it really small and in this sense , it is important for people to like also consider the security aspect , as we do the compression . So this is something I don't think a lot of people are doing , but people are starting to realize that compression may also cause security issues and we need to pay specific attention along this line yeah , I think in general , this is a very important but also very complex issue and it should be approached from different angles , definitely from the regulation , from design and other type of things .

Speaker 4

but talking specifically about device things , I think the very simple thing that we already do in putting some guard rails on these devices so you cannot use them , and some simple things like prompt engineering and so on , you cannot ask a question like how do I design a dirty bomb , type of things , so it's going to be blocked , so it's already happening but definitely needs more attention .

Speaker 1

Okay , no , we have to actually stop it now . We've also exhausted already the planned time in the end to talk to the audience and I have even two more questions that I had to trash , so we need more time next time . Anyway , but let me still summarize just in one sentence . So we all see this or we sense this is a big opportunity and it's already available now . So Gen AI is not necessarily large language model or even small language model . There is a lot of Gen AI techniques that can be applied now with current hardware , in the current ecosystem , and also we'd like to encourage participation in Danilo's working group , which is about Gen AI . So sometimes the calls there's people , sometimes there's more sparse . So think about it . This is the place to move something . Thank you .