Bioprocessing Unfiltered Podcast

Unlocking Smarter Bioprocessing through Better Data Collection, Quality, and Analysis

August 4, 2026

Bioprocessing Unfiltered Ep 10

Everyone wants smarter bioprocessing, yet most teams are still fighting the same battle they fought years ago: inconsistent information, messy process data, and systems that were never designed to support analysis. At this year's Bioprocessing Summit Europe, industry experts gathered on a panel to address these gaps, explore why data challenges persist, and discuss why the hardest part is often data engineering and harmonization. Moderated by Maximilian Krippl, Ph.D., head of bioprocess modeling consulting at Novasign, the panel consisted of Andrea Arsiccio, Ph.D., senior scientist and team lead in silico at Coriolis Pharma; Ayca Cetinkaya, Ph.D., senior scientist at AstraZeneca; Mario P. Pereira, Ph.D., director of technology and business development at ATUM; and Jack Prior, Ph.D., head of process monitoring and data science and AI strategy at Sanofi. Their discussion also explores strategies to transform raw information into actionable insight and practical guidance to make data comparable and usable.


PANELISTS BIO

Andrea Arsiccio, Ph.D., Principal Scientist, Coriolis Pharma 
Andrea Arsiccio is a principal scientist at Coriolis Pharma, where he is implementing innovative modelling concepts for biopharmaceutical formulation and drug product development. He received a master's degree in chemical engineering in 2016 followed by a doctorate in 2020. During his doctorate, he investigated the computational modelling of proteins, with a special focus on the formulation of biopharmaceuticals. He subsequently worked as a postdoctoral scholar at the University of California in Santa Barbara and at Politecnico di Torino, and during this time his research focused on molecular dynamics simulations and bioinformatics. His main interests and expertise are in the fields of modelling, protein folding and aggregation, and freeze-drying.

Ayca Cetinkaya, Ph.D., Senior Scientist, AstraZeneca 
Dr. Ayca Cankorur-Cetinkaya is a senior scientist at AstraZeneca, specializing in upstream process development for biopharmaceutical manufacturing. With over a decade of experience in the optimization and advancement of upstream processes, Dr. Cankorur-Cetinkaya has played a pivotal role in the development and technology transfer of CHO cell line-based platforms for the manufacture of diverse therapeutic modalities under GMP conditions, supporting the progression of early-stage clinical studies. Dr. Cankorur-Cetinkaya earned her doctorate in chemical engineering from Bogazici University in Istanbul, where her research focused on systems biology. She is the author of more than ten peer-reviewed publications and is recognized for her commitment to scientific collaboration and innovation. Passionate about integrating systems biology principles into upstream process development, she leads multidisciplinary teams dedicated to advancing manufacturing strategies in alignment with evolving biotherapeutic technologies.

Mario P. Pereira, Ph.D., Director, Technology & Business Development, ATUM 
Mario Pereira is currently the director of technology and business development at ATUM since March 2023. Prior to this, they worked at Horizon Discovery from 2016 to 2023 in various roles including strategic account manager, field application scientist, and senior scientist. Mario also worked as a principal scientist at FUJIFILM Diosynth Biotechnologies from 2015 to 2016, focusing on developing recombinant CHO cell lines. Mario Pereira completed their biotechnology doctorate at the University of Manchester from 2011 to 2015, where their research focused on genomic environment for recombinant gene expression. Mario Pereira began their education in 2005 at the Universidade de Coimbra, where they obtained a bachelor's degree in biochemistry in 2008. Following this, they pursued a master's degree in biochemistry from the same institution, completing their studies in 2010. Mario then continued their academic journey at The University of Manchester, where they enrolled in a doctorate program in biotechnology. Mario successfully completed their doctorate in 2015.

Jack Prior, Ph.D., Head, Process Monitoring & Data Science & AI Strategy, Sanofi 
Dr. Jack Prior leads MSAT process monitoring and data science/AI strategy at Sanofi, where he spearheads global initiatives in process monitoring and AI-driven yield improvement for biologics manufacturing. Previously as head of MSAT Digital, he led teams working to develop global process data analytics systems and to digitize laboratory operations for Industrial Affairs Specialty Care drug substance and drug product manufacturing. Jack’s career at Sanofi-Genzyme spans nearly 3 decades, where he has led manufacturing science organizations supporting process characterization, modeling, technology transfer, and manufacturing operations for critical therapies treating rare genetic disorders. Throughout his career, he has focused on integrating process modeling and advanced analytics with manufacturing science to enhance biologics production across US and European operations. Dr. Prior holds a doctorate in chemical engineering from MIT and a bachelor's degree in chemical engineering from the University of Connecticut.

MODERATOR BIO

Maximilian Krippl, Ph.D., Head of Bioprocess Modeling Consulting, Novasign GmbH
Maximilian is responsible for Novasign's modeling projects in bioprocessing and helps to accelerate and better understand customers' bioprocesses by using advanced machine learning. He studied technical chemistry at TU Vienna and obtained his doctorate from the University of Natural Resources and Applied Life Sciences.


TRANSCRIPT

Welcome And Panel Setup

Announcement

Welcome to the Bioprocessing Unfiltered Podcast. Each month we host conversations with the researchers and leaders tackling and solving the day-to-day challenges of the bioprocessing industry.

Maximilian Krippl

Hello, everybody, and welcome to today's panel discussion. So we're going to talk about unlocking smarter bioprocesses and whatever we need to do to get there. So talk about quality, we'll talk about data, we'll talk about culture. So we have quite a versatile panel here coming from coming from big pharma, coming from service providers, coming from CROs, all of them working some way or another with data, with with insights into company data, into processes. Not only bioprocesses, even though we're at the bioprocessing summit, but we'll talk also about other aspects of data. So I have all the information about the attendees here, but I let them introduce themselves because I'm sure they themselves because I'm sure they do a better job than me. So Mario, if you want to start.

Who Is On The Panel

Maximilian Krippl

Exactly. We have one microphone for you.

Mario P. Pereira

Thank you. Right. Yeah, I'm Mario. I'm my title is quite quite broad, is Director of Technology and Business Development at ATUM. What does it mean? It means that I do business development and I look for technology as well. And yeah, at Atom we we we use a lot of data to as you I don't know if you had the opportunity to see earlier today, but we do a lot of the use a lot of data for machine learning, but a lot of it is in the upstream of upstream development. So it's more on cell line development, vector design, optimization of algorithms. And that's the insight that I kind of gonna bring a little bit more, but but ultimately if you're looking at at optimizing machine learning methodologies, data is key. I think a lot of the learnings that we have at the early stages, I think can be translatable to some of more the process development side.

Jack Prior

Hi, I'm Jack Prior from Santa Fe. I work in the manufacturing science technology organization focused on commercial stage manufacturing, not the early development that some folks are in. I've worked in process data, manufacturing process data for almost 40 years. If I include grad school at MIT, I got my PhD in in chemical engineering, but the focus or at least the intent of my thesis work was to apply artificial intelligence to bioprocess data about 40 years too early. And since then, I've been focused in you know, either part of my role in manufacturing science or all of my role in on manufacturing and process data. I'll speak tomorrow on the topic as well.

Ayca Cetinkaya

Hi everyone, my name is Ayca Cetinkaya. I'm working as an associated principal scientist in AstraZeneca in upstream bioprocess development team. I'm working on early stage projects generally or a bit later stage as well, doing process optimization and transferring the process for scale uploads. But by training, I'm a chemical engineer and I had my PhD in systems biology. So I'm really keen on applying these systems biology approaches into process development where we we talk about systems biology, we deal with larger data sets, which I think links with the topic here.

Andrea Arsiccio

Andrea, nice to meet all of you. I'm also a chemical engineer by background like Jack and ITEP, so quite similar background here. I work at a CRDO based in Munich called Coriolis Pharma. Actually, we're opening in this very day source site in the US, so we are expanding. And just to give it a bit of background, we work on track product development. We also have analytical services and some manufacturing services. So we have a strong wet lab focus there, especially information development. But me personally, I work on in silico models. So I'm currently leading the in-silico team at Coriolis, working both on data-driven models, machine learning, etc., but also physics-based models. So my background during PhD was more on molecular simulation, especially molecular dynamics. Very happy to be part of this panel. Okay, thank you.

Maximilian Krippl

My name is Maximilian Krippl. I'm head of consulting at Novasign, and we're doing we're a software company doing bioprocess modeling and digital twins. So basically 99% data. So that's that's my background here.

Why Bioprocess Data Is Messy

Maximilian Krippl

Okay, starting out with with the first topic, I want like to cover three main topics that we decided on. One is basically data in general, so messy process data. Is it messy? Do we have enough? Or is it the methodologies that are missing? Secondly, talking a bit more about some modeling aspects. So, what can you do with the data? On the one hand, we know there's a lot of mechanistic models, a lot of AI models, so maybe especially on Andrea Andrea can can add to that, , and also where where is the journey heading? So maybe Jack, you mentioned already 40 years ago, machine learning was already a topic in in that industry, maybe with with some some slightly different aspects as today. One thing that I for sure now can imagine there was less data than today available for for specific processes, for specific aspects. So, my my first question, maybe going first from from big pharma more then more to the to the CRO and service providers. Is it do you see right now more the issue in handling data in getting all the data into a centralized location? You can actually make decisions on it, or or or is it more the tools that that are missing? Or are they missing? Do are we there? Your talk tomorrow is about are we there yet? Digital transformation. So maybe are we there yet?

Jack Prior

It's yes. Boy, that's a wide-ranging question. I don't I could just finish my talk today and then people have a longer lunch. I mean, I mean the kind of the premise, you know, the conclusion I came to and kind of the work that I'm presenting tomorrow is in some ways we are not that much better than we were 20 or 30 or 40 years ago. In some ways, it depends in commercial manufacturing, and it can be better, but it can be worse. And I it's people would assume with paperless manufacturing, with digitization that we're on a steady, exponential upward climb, everything gets faster, better, easier, but it doesn't always because when you have a moderate amount of data and a strong need for it, you type it in. And that's how we started, right, for some of it, right? And then you find ways to get simple pathways in. As you get bigger and bigger, more complex, more globalized, you expect to not type anything in. But the replacement is the need for data engineering, which is a term I don't think I heard before 2019, and then now it's all I hear, right? Getting getting data go from point A to point B when it wasn't designed to do that. If you if you're in the lab and clearly your purpose is to generate data for a very particular purpose, hopefully you probably can get the data flows in place. But if you're if your purpose is to do commercial manufacturing, the systems involved are not necessarily intended to make a data scientist happy six months later who wants to analyze it. So I think the we we have we don't have a lot of data compared to the industry, some of the things that AI is famous for, right? Netflix or Facebook or you know, where there's terabytes of data. We have dozen runs, hundreds of runs. You can only fit, you can only do so much machine learning on a limited number of runs. You only have as many unknowns as you have have have data points to to fit. So I don't know if that's enough of the answer. So I I think it's it's not settled. It's I think there's a lot of pain out there. I mean, it would be good, show of hands. I mean, how many people feel you have all the data you need and everything's smooth? Okay. So we would go next. Yeah.

Maximilian Krippl

And TV, Ayca, you're right.

Ayca Cetinkaya

So I will approach from more early stage process development data. I mean, we have, I mean, maybe more data than we have in academia. In in industry, we are having lots of projects testing those molecules. But I think the data is not as ready as to be directly downloaded from the road as a raw data and usable in for machine learning. There are lots of factors that need to be considered when you are using a data. First of all, is it is there any deviation in this room? Well, the same instruments were used. I mean, they are all recorded in somewhere. In the lab notebooks, it will state if there's a deviation, it will be recorded. But as if you download the data directly from the database, you wouldn't see that information. There might be some missing data points. You wouldn't know why it is missing. Is it like showing as zero? Is it zero? Or is the instrument was not calibrated on that day and it's not a critical parameter. The scientists just say, okay, I'm omitting this day. And how you will fill these gaps while processing this data. It really requires the scientist input to say, okay, this is a good data set, a verification that you can use it further analysis. And it is, I think, really time-taking process. Obviously, it can be accelerated by using AI tools, integrating all these different resources to verify these data sources, but I think at the moment we are not there yet to make it automated.

Andrea Arsiccio

Yeah, so also from my point, I think we do have data, for instance, but it's very fragmented. So if you think about it, we have data maybe in the ELN, in LIMS, in some software from selected instruments. We have data in some spreadsheet, we have data in reports from very old projects. And this creates one problem. The other problem is that the workflows are still not perfectly streamlined, so there is a lot of, for instance, copy-pasting, then transfer into Excel spreadsheets. And there are also errors into this copy pasting. So the data is heavily not homogeneous, and these create some problems then when you try to process this data via machine learning algorithms. What also is hard, especially when you have a lot of different use cases, is to align on the namings, on the nomenclature, on what we call ontology. So you need to have this clear ontology if you want to apply then data-driven models. And it's kind of hard to agree on that on a clear nomenclature that works for all the use cases. Then there is this mixed miscommunication between the data scientists and the end users. And the end users sometimes perceive the new protocols as just adding more effort because maybe they need to duplicate their work, and then that option is low. So a lot of problems, both in the let's say, distribution of the data, so they are not centralized, as well in the communication, I would say, between data scientists and people working in the lab.

Mario P. Pereira

Right. Yeah, I again as I said, , my perspective is slightly different, but but again, the data, the data part it's it's maintained. So one of the things that I'm lucky to to have is when I joined ATUM, ATUM was already 20 years of of thinking about this type of issues. And, and so the foundation of the company was kind of a limb system that was internally developed. So the company is based in San Francisco, and one of the things that there is a lot in San Francisco is people programming. So and and so they took advantage of it, and rather than trying to fit different LIM systems and try to put them together, design it from the beginning. And what that helped with it helped having a limb system that evolved with the company to collect the data that that the for the company needs. So we it started with with controlling the data first. So for the protein engineering, vector design, cell line development, how you design the the data sets that you need and how do you generate that data that you need. And at the same time as you create cell lines, for example, and now we move more to towards cell line development as a company, we are starting to generate the limbs that we need to collect that metabolic data that we believe that is useful for the next level. So we are now it's so the foundation is is like of making sure that the at least for the way we see it, it making sure that the data is as clean as possible and control it as much as possible. I think it's probably quite trivial what I'm just saying, but but one thing is people talking, talking, but then it's hard to put it together when you start it from the beginning, it really, really helps, and I think that's helping quite quite a bit to us.

Maximilian Krippl

Okay. I forgot to mention I will also give you the options to ask questions throughout the present throughout the discussion, not only at the end. So I will continue, but afterwards, if you have anything else to ask, we'll we'll have time for that as well. One thing that was mentioned was data access. So, I mean, if you now with the technology that is available now that for data storing, data lakes, ways to model, ways to visualize, maybe that comes even earlier than modeling data. I mean, starting now, you could probably potentially implement it quite well with the tools. But there's a lot of data, especially now talk looking at at Jack and Aisha here, that historic data, right? That potentially it's really difficult to access. It's there, but it's in a format that is not supported. It's a lot of it, it's a lot of hassle to to get it. Are there efforts in trying to get this data or do you or is it easier to just say, okay, from this point on, we're just doing it right and then and then basically use this this information? Can you give us a bit of an energy on this?

Ayca Cetinkaya

I think we can completely ignore, okay, if we just start from here. We need to make use of this data as much as possible. But I think it is more individual efforts, like we are really putting in the time and as much as we can go to the maybe not that far back, but more recent years, and if the scientists that run this experiment are actually still in the company as well, to just really get quick answers. But I think one advantage maybe as we are moving more complex

Context, Deviations, And Fragmented Systems

Ayca Cetinkaya

molecules, because most of the historical data is with standard maps anyway. And but for for the starts, as we deal with more complex molecules, for these molecules we can make it right. And definitely that that there's a recognition of how important it is and adapt the systems accordingly.

Jack Prior

Yeah, I mean, I think generally in commercial manufacturing, we're pr often pretty good at keeping you know, accumulating data and not starting fresh. But it is a challenge because we're, you know, at the layer of process data scientist, you're looking at six different source systems at a time. You got to deal with LIMS system, MES system, data historian, ERP, genealogy, et cetera. And those aren't fixed forever, right? If you have a company that has custom system and it stays the same, you know, we tend to change those underlying systems frequently. And again, as I said, they would none of you know, some of them are definitely not designed to feed me. I don't know why I thought there would be 20 years ago when they promised that they were coming with paperless systems, but you know, you have to build the new interface and then often have to find create a repository for that legacy system because it it will have some kind of GMP archiving, but it won't be readily queriable. So you need to manage this kind of data that you used to query live. Now you keep it separately. And I think you know, it's so it's it's a challenge, but in a way, if you had it sorted before, you probably are better off than it's the new systems that you have to readapt to, right? Something changes, now you have to build a new interface and the mapping of the language of the terms, right? You what you know, we used to call it bioreactor pH, now it's reactor pH, or it's product name, you know, that the the strings can be tough to map up because ultimately you want a nice rectangular data set that you can pop into jump or some kind of analysis, but but the biggest challenge can be just harmonizing names over time across different manufacturing sites, across companies. A lot of things can get in the way of the sort of the the final experience of being able to do data analysis and machine learning.

Maximilian Krippl

Are there questions from the audience? Sam, right?

Audience Member #1

Thank you. It's very interesting to listen to the discussion. I have maybe to that point a question because that what you're basically telling me is that you have kind of messy data in terms of you have different labels, whatever. So you're part of bigger companies, , I would say most of you. So are you try having standards inside your company that you're kind of in ways you sh share data between divisions, between scientists, or is it is maybe that kind of the the problem that you I mean definitely there are standards, but the reason why sometimes the data is massive is just not because we don't have the ready systems, but because of the nature of what happens in the lab.

Ayca Cetinkaya

And also the technologies are evolving, so we are using different technologies to measure. I mean, I I think it's very random now. For example, for measurement of the cell densities, now XR is retiring and industry switching to Bullo, and they are not using the measuring exactly the same, so you will need some correction factors in between. But at the time you were using XR, you didn't need to record that in a specific database. It is XR because it was coming from that, it will be obvious. But now when you compare historical data and the current data, you can't just combine this data, assuming they are the same, you need to have these considerations. And obviously, these are all reported. It needs to be GLS compliance, but not in the same database as it is. So it is this information, I think it's the same problem. It is all different places, how to capture all this data together. I think this is the main challenge.

Audience Member #1

Thank you.

Maximilian Krippl

Other questions? Okay, then what I heard from you, Jack, before was that you basically said there is there are a lot of tools, a lot of things that must be done is basically connecting them. May it be selecting the right strings or may it be making the the electronic lab notebook work for you. Maybe a question this direction. Do you think the tools are generally the question to all of you, do you think that the tools to work with are generally all there and it's just about connecting them for your purpose? Or is there something that you're missing in your daily work that that you would have to develop by yourself or to custom develop that that makes post that would enable you to make more use of this of all of this information?

Jack Prior

Yeah, I mean I think the tools exist in the world. There's several vendors here presenting you know powerful tools. As you get to larger and larger companies and global companies, there's a strong bias to standards to converge to single tools, and ideally a single tool that covers a wide space, right? Like, you know, if you're senior management, you'd you know you don't want 5,000 tools in your company, you want less. And so when you have a big company like SAP that can handle hundreds of different business needs, that's perfect. You drop in workday for HR or something like that. So I think there is the challenge is I don't know that we have a good universal tool that covers all the things that would be of interest to us. And and when we have a preference for individual things that's seen as kind of non-invetted here, people ever have all, you know, people would want jump, mini-tab, SAS, right? In some cases, there are kind of apples to apples tools. In other cases, they're apples and oranges. And it's our, I think our challenge is to bring the the right tools in for the right purpose. That often involves making the case to decision makers that I do need this tool and that tool. They're not the same. You can't so I don't want to throw any particular tools under the bus. But sometimes it's just not the right tool for the need when you come down to doing the work with it. From 10,000 feet, it looks like it might be the right tool. So I think if you have infinite budget and ability to download or purchase there, you would create a very nice set of tools to align globally in this space, it's a little bit more challenging, right? We don't have, you know, and we and honestly, we'd need two or three universal tools to really compete to get the right use user requirements. I think the challenge for us is we haven't always defined the user requirements clear. What do we, what would an adequate tool need? And it's it varies. I'd also say that when you're I'm I'm sort of passionate about root cause analysis and troubleshooting, you can't always anticipate what you need to do the troubleshooting until you have to go and find some fact or some set of data. And we, as an industry, especially in commercial manufacturing, we like to validate everything. We like to have pre-tested it for all purposes. And that's hard to do, certainly to the last mile or the last kilometer of of work. You know, you're you're gonna have to do some analysis you didn't plan on because you didn't plan on them problem.

Ayca Cetinkaya

Just maybe I will add one thing on that when you say about the root cause analysis. I think the type of the data we're analyzing also depends on which tool we are using. For example, when I think of in terms of the bioreactor data, we have both offline data we measure regularly, like every day. Let's say we have samples from offline data, but also the bioreactor is saturating online data and how much we leverage from that data. And the tools that we can use for analyzing offline and just combining this offline data and online data together is another challenge because the dimensionals is quite different. And to specifically reach out this reactor softwares and getting out of it everything and analyzing together with data. I'm I'm sure there are tools, but I think it's not established across the all the company to use. In the same same way.

Andrea Arsiccio

For me, just a short comment. So I agree with Jack. So I think there are a lot of tools there already available. So I don't think that's the main limitation. I think the main struggle, at least in my opinion, is that it's still hard to find a clear agreement on the format in which you want to have the data, all the data, right? So what happens frequently to me is that I have the idea of a model in mind, I extract the data, then while training the model, I realize, you know what, maybe I would have also liked to have that other piece of information for this data. And then you need to go back and add this type of information, this metadata, and so on. And that takes a lot of time. So I think what really is the hardest part is agreeing on the data foundation, so how to like harmonize the data and in which format to have them and the data infrastructure, so where to have the data as well. So that's the most challenging part, at least in my opinion.

Mario P. Pereira

I don't think I have much to have. I think my my my point of

Saving Legacy Data That Still Matters

Mario P. Pereira

view is tools are out there. You can data, it's you can collect as much data as you as as you want with all the tools. I think it's a lot of it is about how you connect them and how it fits what you need. What company A requires, it's not what Company B requires. So of course it's important that we all share to understand what we can take, take a message from all. But but for example, again, I'm going back to the experience that we we have. We we we we are about 160 people, and 10% of the company is just developers developing internal LIMS because we feel really, really strongly that we need to be the ones developing the connection points of the different tools that we have. So it's 10% of a company, it's quite a lot just on the developers when our main purpose is not developing software. Our main purpose is to make sure that we have really good cells, really good proteins, really good, really good DNA coming out. But we really feel that for us to deliver on those three points, we have to invest on that. Make sure that whatever tools we bring in and when we update, we keep them coming. Again, I'm not saying it's gonna work for everyone, but it's it's that's at least helps sharing that. And I think yeah, bringing the tools is important, making sure that they are connected for your needs. I think that's based on my experience, is kind of the one of the most important parts.

Maximilian Krippl

Okay. Maybe you said a good point. We can collect all the data in the end. We need to make some some conclusions from this. Just collecting data doesn't help us. So I want to switch gears a bit. And basically, everything we can do with data, I just call it modeling. Modeling can be anything, but when we talk about processes, we can do models on on bioprocesses, for example, about there's different types, like going from very mechanistic ones to very AI-driven ones, but also in terms of now, since a couple of years with LLMs, you can potentially use entire text, you can use reports as in as as input to your model to generate maybe new reports for regulatory or for clinical. My question here is in the last couple of years, what made for your organization or your personal work a big difference? And maybe also what what didn't in in in using such models? Like do you use them? If yes, which of which and what what actually makes a big the biggest difference?

Andrea Arsiccio

Sorry, or I guess hard. Yeah, actually, so that's a good question. I think we use basically all types of models, so mechanistic models, more hybrid models, and purely data-driven models. And I would say it really depends on the focus, right? So obviously, if the system is not too complex, I would then really go for mechanistic modeling. So as a chemical engineering, as a chemical engineer, mechanistic modeling is a like my it's my heart, so it's my preferred option because it can extrapolate, you have less hallucination. It's also easy for verification and validation activities because it's very interpretable, so it's much easier also to propose a purely mechanistic model to regulators. At the same time, for very complex systems, it's very hard to have a purely mechanistic model. And that's where, for instance, we use a lot data-driven approaches as well, trying to avoid complete black boxes approaches where you have no clue how your machine learning model works and you lose any idea of how to interpret your data. So that's we try to avoid anyway. So, just to give like my my perspective, I've I would never consider actually using a purely data-driven models, for instance, for digital twins for online process control, because the tendency to hallucinate is still too too high. And I would not even try to produce a VNV document of a data-driven model for online process control. That would be a nightmare. So let's say, again, it's a combination of different considerations we have to keep in mind. Hybrid models have combined the benefit, I think, of both words. So at the moment they are they they are very popular, and I think they have a lot of advantages in the end. The main disadvantage of hybrid models is that you need to have both domain expertise as well as data science expertise. So sometimes it's hard to have people that combine this type of knowledge. So that's that's the main the main thing. But coming back to our like work at Coriolis, for instance, so we use data-driven models, purely data-driven models for, for instance, liability identification. They are low-risk models, so there we are safe with purely data-driven approaches. When we go into process optimization, there we try to use more mechanistic models, for instance.

Ayca Cetinkaya

Personally, I mean, as I presented earlier this morning, my main area is on metabolic modeling, so it's more mechanistic in that sense. But I think as we see in the talks from the session, really we need to divert towards hybrid modeling and really need to understand what are the limitations of different approaches. I mean, especially for the metabolic modeling, for example, it's purely based on metabolism, so we wouldn't predict any restrictions on this productivity based on the post-translation modifications from there. We need another modeling approach to fill that gap. Or similar, when we are analyzing all these large data sets we are generating from model predictions, all these flux distributions. Again, we need the data science to be able to predict all these flux distributions, and we can learn from those to feed into digital TIVIN or a machine learning approach. I think it's more diverting as we have different modeling approaches, combining these, just filling the gaps for their limitations will be the feature.

Jack Prior

Thank you. Well, it's a broad topic models. I mean, I would say mechanistic models are helpful and I think more and more possible, like even your talk or some of the talks I've seen with you know, with genome mapping and the tools to understand what's happening in the cell, where I would have dismissed it 30 years ago that you miss a few equal reactions and everything is a mess. But now we know a lot more of the reactions, so there's more possibility there. I think mechanistic modeling kind of relates to hybrid modeling. When I did the first modeling early in my career, I developed a kind of a digital twin in probably 1999. It was to understand for the dynamics I was seeing in the process, how much could be explained by the basic, you know, my mechanistic understanding or empirical understanding of what celliculture was doing and then what's the delta. At the time, there wasn't the idea of hybrid modeling to learn from that delta. When I talked to my experts, even in anticipation of this panel, that we're having some success with hybrid modeling. It makes sense. For AI, I think it's in manufacturing, it's often overhyped. I think we jumped to MVA quickly because they're very convenient, well-marketed packages on the market. We don't, there's not good neural network packages, ironically, that people might use, but there's a lot of MBA and the data is kind of great because you can go to a conference and it's unintelligible, so there's no IP risk. But I personally like XY plots, right? That's my machine learning is if there's something wrong in your process that changed or that is changing and causing an effect, then show it to me as an XY plot. Don't show me a SHAP plot. Don't say there are 20 important parameters, is usually one. If you're lucky that you can measure and trace for root cause or for root yield improvement. And when you start telling me there's four or five or six, very quickly you're overfitting the data. And so that's just my take on sort of what I would call AI 1.0, AI 2.0 with LLMs

Tools Exist, Requirements Are Fuzzy

Jack Prior

and agentic, whole nother story. I'll talk about that at the end of my talk tomorrow.

Mario P. Pereira

I have nothing to add on this one because the way we use modeling is for completely thick for things, protein engineering and and and gene and common optimization. But one thing that I can agree is on the XY plots. I've I'm a very simplistic person. That's where we try to do when we make those models and how to test them. And , and that's that's something that I I completely completely agree. Yeah.

Announcement

Are you enjoying the conversation? We'd love to hear from you. Please subscribe to the podcast and give us a rating. It helps other people find and join the conversation. If you've got speaker or topic ideas, we'd love to hear those too. You can send them in a podcast review.

Maximilian Krippl

Are there questions from the audience?

Audience Member #2

Hi there, I'm Joe Egan from Teaside University in the UK. Yeah, a lot of talk about hybrid modeling. I'm interested to know when you're sort of fusing your mechanistic models with your data-driven models, is there a kind of a unified framework, a best practice to do that? Because I imagine there are various ways you can bring these together. And I just wonder from your combined experience when you've been developing these, have you tried different approaches to to form these hybrid models and kind of got to a bit where you've you kind of fully understand the best way to fuse these things together?

Maximilian Krippl

I'm a bit biased because I work for a company that sells such a solution. So well so generally the what at least we are trying to do is to whether we use a mechanistic, a hybrid, or a machine learning or an purely AI-driven model, that's basically I think there is no right or wrong, right? It it depends on the task, and therefore the solution should reflect that. So so the the the approach you're using and and the ideal the software and the and the the infrastructure use should reflect that. So you should have like full flexibility to go completely in one direction. Maybe also for a proof of concept, you want to start easy and you don't or simply you don't care about the accuracy of the model too much. You want to know whether what you expect is actually there, and then start collecting more data, start adding more and more the aspects of unknowns to it, and then shift towards towards a combination of ideally a combination of the mechanistic part and the and the and the machine learning part there. I mean I come from process, from bioprocessing, so it's it also depends a bit on the units.

Maximilian Krippl

So I think there are definitely process units that are more where you can measure more simply, and you can there are more theories out there that that you can apply, and others that are where it's really difficult to actually make make proper assumptions other than I don't know, cell growth and does something, just saying because you you you would then need to add a lot of measurements to to that to actually get insights or more more insights. And I think it's it should enable for going from a proof of concept to a global feasibility to a feasibility for a specific project, maybe to multiple projects, and maybe at some point global if you if you get there. But yeah, it's it's unrealistic to say hey we have to buy all of these sensors and start with a full-blown approach because at least this is for us it doesn't work like this. Yeah. Any other questions? Maybe one question in this regard. What what we like to do when we when we start collecting data or modeling a process, we we normally like to get all the data we can, all the data that is recorded. But most of the time we wind up using only a fraction of it because it's the things that are actually controlled, the things that are actually first of all measured and secondly controlled. So what what we see from from how we apply it is that there is way more data generated than what we currently actually actually use.

Maximilian Krippl

And this depending on on the project, this can be quite a huge factor. Sometimes it's we're just using 10% of the data because that all all the information is in there. Do you think there is an overhead on the data that we're actually collecting? And that is there it is there a big part that we can actually never really infer any information from, or would you just say if I can record it, I record it, I measure it. I don't know if I ever will will ever use it. Maybe it's just gigabytes and terabytes of data that I will have to maintain and and I will never use it. Is there what's your opinion on this? Or would you just buy more hard drives?

Mario P. Pereira

I can even start on this. I think it's I think collecting it i it's worth it because you never know. That's the easiest the easiest answer. It's what happens is that we'll never never never know. But but I'll give you I'll give you the the example again, going back to what our experience is we've been collecting data, and as the company evolves, the data that we collect evolves as well. So like start DNA synthesis, starting protein engineering, move to vector designs, move to cell line development, and now with cell line development, we have done more than 200-300 projects. Everything, every time we do a Fed batch, we collect metabolic data, right? And , and that was recorded in a specific way. That metabolic data was primarily used to analyze each project at each time, but now and that it was but now we have a like a really nice source of data that we've just been keeping it just because you never know. And we never know could come very soon, where if the company decides to progress to something else, then it started to have a kind of a set of 200 different projects or 300 different projects with cool data collected in a very similar way, with with cell lines that were in the control environment and the data that is also collected not in the same way, but could potentially be collected interacted.

Mario P. Pereira

And so you started. I think it's important, even though if you're not going to use it right now. Problem is now how you can keep it and the the environmental cost, all of those things, but but as it stands, I think there's a lot of value on it because as also the ways machine learning and actually AI works in terms of trying to connect dots together, that it's hard for us to do it. It's it could could potentially help. So, my opinion, yeah, it's almost like an insurance policy, and it also increases the value of your company. I think it's quite well, hopefully it sounds quite obvious, but but I think that's that's that's where we stand at the moment. Yeah.

Jack Prior

Okay. Anyone else want to go first? Or I'll jump in. All right. Yeah, there's a couple of, I guess I can tell one story from my early career. When I first started, it was Genzyme's Alston plant in Cambridge, Boston area, was the most automated plant in the world at a time. This is 1997 or so. And we collected 500 parameters on our data historian, pH, temperature, levels. And

Mechanistic Vs Hybrid Vs AI Models

Jack Prior

what I found from a root-cause troubleshooting monitoring point of view, there was always one more that we didn't have that I needed. I need the temperature in the buffer tank. I need the level in the media tank, you know. And so we would grow to you know 50 points every few months. So we would add those in reactively. When we went to upgrade our historian and put a new one in, I said to my head of automation or my partner, I said, you know, hard drives are getting cheaper every year, right? There's a curve, right? And when this process was designed, I think a gigabyte of space was a million dollars, right? But at that moment, a gigabyte of space was, I don't know, it might have been 250, you know, 25,000. But I said, what if we measured everything in the plant, everything that can move, the PID tuning constants in every loop? Because when something would go out of control on them in the morning, you would wonder, did the operator play with the constants last night? We weren't perfectly locked down in terms of admin or or the or the automation people might have intervened. So we went to 25,000 points and I was very happy. So I that was good.

Jack Prior

But I think more recently, sometimes we there's kind of empty calories of data that can distract from the core. You know, you in some cases we'll do very rigorous data reviews of 246 parameters, maybe at the expense of really focusing on the few key ones, or we'll collect gigabytes of data, thousands of data points, but we'll lack a key, a key point, right? Like there's a dozen points that I need. I'm ready to fly over the across the ocean to type them in. Right. So you I, you know, it's kind of again, I'm revealing my talk tomorrow, which involves a scorecard that tries to really quantify whether you have the the right things in the right way. So I think it's a mix. I mean, it's I definitely agree with it's great to have it if you and just in case, but don't miss the basics, right? It can very easily you can feel good about quantity when you're missing quality.

Mario P. Pereira

I agree. But it is for me more about you also need to have some level of pragmatic approach that if you miss it, you miss it. You your you your decision point at the moment is as best as the technology we have at the moment. So it's kind of having that that kind of approach and understanding that, oh yeah, it's me. Well, it's it's the best thing we could have at this given moment and be okay with with that. So it helps, it helps on on yeah.

Jack Prior

Maybe I just make an example. You mentioned like managing online data earlier. You know, the one key thing you need between online data and offline data to or discrete data is some mapping, some feature engineering, some feature extraction. That's one of the six things I talk about. And so in one case, we have online data, but we don't take a number out of it to the offline to the discrete data. We I need the pH once a day. You know, I it's great that I have it once a second, but I need the once a day number to come into my daily data or the final value or the beginning value. It's a trivial little thing, but it's very easy to miss as you start getting a broad organization and you start having those decisions and engineering happening with the people who don't, you know, maybe aren't on top of it. So it's it's a good question. It's it's it's squishy.

Ayca Cetinkaya

Yeah, I think we I mean we need to keep having more data. I mean, that's side as well. And I think we also need to consider how we leverage from that data. Yeah, we collect lots of data, but for most of it, we don't just look at even the data. Is it just recorded automatically, especially in the in the case of online data? And but there's some hidden information there. We need to really concentrate on how we can get learn from it more. And with this new AI approaches and these things, we can leverage more from those. And as earlier, we actually in the talks, we see, for example, the span medium analysis. I mean, as it and it it is very simple, right? You see what is limiting there. If you don't get that data, you wouldn't know. It's not as something maybe with the the probes that you in your in your talk that you mentioned now it's possible to get this online and monitor them regularly. But if we think of this offline sampling and analysis, if we don't intend to take this data, then we wouldn't know the root cause, real root cause, and how we can improve further. So definitely we need to we need to have we need more data.

Andrea Arsiccio

Yeah, so I have a personal problem discarding things, so you should have a look at my basement, you would understand. So in summary, yes, I for me we should definitely store the data, and I have some examples about it about this. So I work a lot on molecular dynamics simulations. Those are the very bad guys, so they generate massive amounts of data. You can generate terabytes with a single run if you really want to store all the coordinates of your molecules and so on. And the way we use molecular dynamics simulations, we'll talk more about that in my talk tomorrow, is to study the interaction, for instance, between proteins and excipients to guide the selection of formulations. And one could think, yeah, why not? I could save only the coordinates of the protein, and then it's the protein I care about for analysis. But that's sometimes a blind approach because maybe then you need to also study the interaction between the protein and your excipients, and then you also need to save the excepient coordinates. So if you already save them from the beginning, you save a lot of effort in the end. So it's true that it's very expensive to store the data, but it's only more expensive to go back and generate the data again. So my suggestion is keep the data.

Mario P. Pereira

I can also have a sarcastic sarcastic or not pragmatic or pragmatic approach is well, I think it's important to keep the data because if the more data you keep, the more spanners you break you throw to the AI, and everyone says AI is gonna remove jobs. So that's the way of keeping jobs. Because if we have more data, collecting more data, throwing more to those machine learnings, make sure that it's so confusing that it's always spinning, always spinning, always spinning, and we keep going.

Maximilian Krippl

So that's that's another way of are there maybe not questions but comments from the audience? Do you have any anything to add to this. Are you more on the more data side? Maybe raise a hands who is more on the wants more data? Okay, so everybody else has enough or okay. Good then maybe coming to to the last point which is how to actually talk about data collection, about modeling, about getting getting insights into this data. Normally they start always with a certain proof of concept, right? You're not doing this for all customers or for the entire company globally. You're just starting out with specific aspects talking about or considering what we just talked about data acquisition, getting insights, visualizations, ideally Y and Y and X axis, or maybe more complex models. Can you explain a bit maybe some some success stories on the one hand that you saw over the years where something that grew just from a small project into something that was enrolled for an entire department or maybe globally worked or maybe also some initiatives sorry maybe also some initiatives that that failed and that now maybe you you know better on how to do it some or to to educate the the attendees here on a bit on on how to do it better.

Andrea Arsiccio

Open Go for you okay good so I think it's easy to start from things that succeeded so yes I will start from that. So I think for success it is important to consider two aspects on the one it's important to consider the end users and where they see a real benefit from addition from digitalization and so on. So even simple things tools for visualization so that people do not have to collect data in the lab, then transfer like manually into an Excel file and then produce these XY plots but maybe they have automatic parsers that go from the instrument into a software that immediately visualizes the data and does some statistics on the data. That's simple but scientists immediately see a benefit on that and we implemented that and that's the that's that's not hard but that's already a big step and people see a strong benefit in it. On the other hand when I think about mechanistic models and so on also there are for success I think the important is that you focus for instance on those unit operations that are a real business case so where you can easily track KPIs and you can convince even the management like you see it's what's really worth implementing this model now. So we like we worked for instance at Coriolis model for LIOFILISION to like streamline the optimization of primary drawing for LIOFILISION trucks. That is one of these unit operations that costs a lot of money and you spend a lot of time also in running a LIO cycle so if you can reduce this time or achieve optimal conditions quicker it's it brings real strong benefits. So these are cases where modeling is really helpful and brings to like to success in the end. Then failures yeah sometimes I I think as you were mentioning things tend to fail when for instance you don't have a real business case or you aim too high. So I want to build now an end-to-end digital approach and then without having clear data foundation that infrastructure nothing then that for sure brings to to failure or when you don't have KPIs so you you work on a on a case study that is not meaningful in the end for the company maybe you spend a lot of time modeling something that brings no value.

Ayca Cetinkaya

So these are cases where in my opinion like it's not worth to add a model in the end I think it takes some time for convincing people who are not familiar with these modeling approaches to say actually it is working. And you definitely need to show that actually you need some do some verification on it then because if you just go especially I mean our company I was in was probably in the same big companies you have okay I have this tool very cool this predicting and I'm I have this hypothesis and unless you prove it then we will say okay but so what do you need to show that it's working I think and as you said you spend time on it to establish these protocols these methods and actually to get approval for you to spend time on it you need to show that actually you need a proof of concept to show that actually it's working and it's a challenge because if you don't have these verification experiments you can't prove anything but to do this verification experiment you need it will ask even more time to do that. So yeah I think the those results some of the results I have shown was a part of our my proof of concept and it takes time but I mean not all the model predictions work. Sometimes you have predictions and no it doesn't come through and I think this is normal to expect otherwise if then it will mean that we really mimic what's happening we have the perfect model and predicting everything. It will take some iterations and we need to learn from why model is not predicting and adapt our model

More Data Helps Until Quality Slips

Ayca Cetinkaya

but I think it's just change of mindset not expect know what to expect actually it is it's really important.

Jack Prior

Again for really thought provoking question to think about sort of what what succeeds and what fails. I mean I think reflecting on it I think the key to digital transformation is that the person the end user as you say has to have a positive experience. You know when we rolled out SharePoint there was no need to do change management right people just well this is better than having a hard drive when we rolled out teams I would say generally pretty positive people just use it you know workday I think yeah I mean so that when it's an immediate friendly experience to the user better than before then that tends to to grow. And then I think you have to have if it's a project it has to benefit multiple stakeholders. My strategy early in my career I actually have a deck from 2003 on this is like it's like a Tom Sawyer approach that's a US analogy but like if if you can if you're a data scientist and can create a system that makes the plant manager get graphs in the morning that helps him do his work, you will get data science data. If you come to a site and say type in data you didn't type in before or type it in twice I it'll be useful to me in two weeks and you won't really know whether what I did with it.

Jack Prior

So I mean I had a meeting on Friday before I came here one hour interview by customer experience, user experience experts of talking about of digital twin project we're working on. It's like revolutionary to talk about what is the purpose of the twin we're going to make and what should the how are you going to use it. Right. I mean I most people want to do twins because they are really cool you know thinking about how you use it. But I I'd say the most important thing is is it a good experience? I also like to say that you can overgovern even pilots and you can over govern standards. I like to say you can't have a best practice until you have two practices and know which one is better. Sometimes we try to jump to the best practice without actually having sort of a Skunkworks competition of different ways to doing it skunkworks projects can't work forever right they can't scale to the to everywhere but those are the things that come to mind as you ask the question.

Mario P. Pereira

Yeah I don't think I have much much to add to be honest at this point. I'm banking again when once you once you have one tool you you everything once you have an MR everything is enabled right no but really the only positive I've seen I joined I joined Atom three and a half years ago three three years three and a half years ago and and this this limbs kind of actually changed my mindset in terms of how how you should look at limbs the fact that it was done from the beginning with with the mindset of the company philosophy really helped and I'll tell you one really simple case study I'm I have a lot of background experience in cell line development and the first time I joined the company and I went to the first cell line development meeting I've first this were just using this software and I was like wow this is quite complex and complicated but but once you started to understand what the rationale of developing those parts and the and the and how easy it was to make decisions it re right you know in an hour meeting we were looking at like 12 projects at the same time and making decisions on how to move those ones in the next two weeks. It started in the second third week I was like okay now totally understand this. So again just it really helped the fact that the things were designed and collected in a way that it makes sense for the company I'm not saying that what we works for us will work for everyone else but but at least something that was developed to work for for for us then I think that's kind of the the way to to look into it.

Maximilian Krippl

Okay. Thank you. Are there comments or questions?

Audience Member #3

Hi as someone is newer to machine learning before I was always in the lab just working on analytics. So right now it's a lot of just cleaning up data and working with systems, fixing the YLN systems and stuff like that. When do I get to the modeling? No but seriously like what's what's the workload like like how much is it just dealing with all the data and systems versus you actually get to the modeling.

Maximilian Krippl

Well this is for me this is a question of of of scope. So you can potentially quite quickly figure out whether trying to to model a certain data set or trying to model certain aspects makes sense. It's you don't have to have all the data in hand at the beginning. What we normally do with with our customers is okay as a feasibility send us just the last couple of runs send us the data you have at the moment is it online is it offline and we just see basically XY maybe xyz plots to see does this does this actually reflect what I what I want to my model to to do. And then from there I think you gradually increase. So normally I think at the beginning it took us it took when we worked with customers a couple weeks until we got the first data sets but with the right expectation management you can really when you start directly then they know what they have to provide you you know what you need then then you can you can start after a couple of days so that's and then how far you want to go that depends then on the on the success or or on the on the ROI or on on how much people need those those insights to perform the work. No are there more questions or comments so maybe just a comment or a reflection so I think the discussion is a really really interesting and it is a recurrent discussion in fact because we know that a model is practically the soul of engineering so we have engineering because we use models and we have to acknowledge the fact that our business is probably very peculiar because it is very difficult to have a working model with at reasonable costs and we have to acknowledge that. So I do think that when we have a platform it makes a lot of sense to invest in modeling so it's probably situation that makes more sense to invest in modeling. And again we have to acknowledge the fact

Pilots That Scale And Projects That Fail

Maximilian Krippl

that we will have to invest in it. So it's not something that is straightforward to get and we will have some time and some investment in order to have a a platform model. A platform model for a platform. So this is probably situation where modeling is more profitable in our business if we go away from a platform then the benefit cost ratio is probably no longer that straightforward to see and I think yeah so so we have to rationalize these ideas in order to have the right investments in the end.

Mario P. Pereira

So this was my reflection at this moment out of the blue thank you for that can I just again as I said I'm not from this field one thing that's every time I come to to this type of discussions everyone talks about models models models and and and to answer your question or to talk about of the of those topics when you get to model I think most important is what you want to model right and if you want to look into the data what data you need to develop that model because if you want to make like an optimal the holy grail you're gonna have to start to clean a lot and you start to do a lot but if you want to like you can look at the problem as the whale or you can look at the problem as okay this is the specific specific part of of of the an organ of the animal right so it's the heart of the animal or whatever kind of a bad analogy but I think people are getting right it's and and and and it goes again again on on on kind of Rui was saying is is kind of it's those things are expensive. So if if we can start looking to these things on a more scale type of thing you start to look into okay I'm for this model I just need this data. And even though I have all the data that I want and just in case you don't need to use it at all the time. You cannot try not to be overwhelmed from it. And this is for me being quite a pragmatic type of approach because sometimes I get really overwhelmed when I see some talk about models especially when I see a lot of X and Y's and Rick letters and so on.

Maximilian Krippl

That's why I really really agree I never thought about it but I really agree the XY way of looking at it it's much much much more simplistic yeah okay I think those were our concluding remarks so I want to thank the panelists I want to thank the attendees just a couple words I should give her talk today in the morning so if you didn't see it maybe and and if you're interested in the topic go talk to her directly but all the others still have the presentation so Jack will be presenting tomorrow shortly before lunch and unfortunately Andrea and Mario are going to present at the same time tomorrow so you have to decide whether you go through one or

Audience Questions And Closing Thoughts

Maximilian Krippl

the other with this we will not talk about models so if you are interested in models feel free go to them thank you

Stay Connected

Follow us on Spotify