Better AI Catalyst Data Speeds Up Carbon to Fuel Shift

In a SLAC-led study, four laboratories collaborated to ensure reproducibility in AI training data to predict new catalysts for future fuel production.

To turn abundant carbon dioxide into valuable fuel, we need a fast and efficient way of determining which catalysts work best over the longest time. AI models have the potential to help guide catalyst selection, but as with internet chatbots, AI models are only as good as the data you put into them.

By convening four laboratories from across the nation to test an experimental carbon monoxide-producing catalyst, a key first step in turning carbon dioxide into fuels, researchers at SLAC National Accelerator Laboratory have demonstrated the importance of generating highly reproducible experimental data when building AI models for investigations in science. They published the results in Nature Catalysis.

Our findings are a reminder to exercise caution about what information we feed into a machine-learning model, and how the consistency of experimental data can influence the reliability of the outcomes,” said Selin Bac, a postdoctoral researcher at the University of California, Santa Barbara, and first author on the study.

Speeding Up Catalyst Development – With AI

With a good AI model, researchers can enter conditions such as temperature, length of time of the reaction, and catalyst formulation, then run the simulation and see a prediction of how well the catalyst performs. They can then confirm the predictions with a few well-designed experiments, ultimately speeding up catalyst discovery and implementation at a global scale.

In addition to saving time and money, such models can also explore conditions that are difficult to achieve in the lab. Most lab catalysis studies can only look at short time periods (days), but catalyst deactivation occurs over time (months to years) due to buildup of impurities and repeated exposure to high temperatures.

AI models need large amounts of high-quality data for training. To generate the data, the four labs performed a set of round-robin experiments, in which multiple laboratories conduct the same tests to evaluate reproducibility using previously agreed upon protocols and the same rhodium-based catalyst.

Squaring the Data From Round-Robin Experiments

To the researchers’ surprise, achieving the same results from four labs working independently was harder than anticipated. When they got together to share their results, they realized that they had a problem. 

Each of the four research teams produced results that varied in the amounts of carbon monoxide and methane, an undesirable side product, produced. The computer would not be able to learn from four sets of data that contain different outcomes.

It was a bit of an eye-opener,” said SLAC staff scientist Adam Hoffman, senior author of the study. “This experience shines light on the practical challenges of including real-world data into machine learning models.”

Painstakingly the teams evaluated their methods. Through rigorous testing they found a handful of sources of mismatch, with one of the biggest contributors to the variability coming down to how hard the mixture was shaken or stirred. 

With further standardization across the four labs – which in addition to SLAC included groups at Pennsylvania State University, Stanford University and University of California, Santa Barbara – the results began to look more consistent. The team outlined several recommendations to strengthen experimental reproducibility, including enhancing the consistency of reactor design, operating protocols and experimental conditions.

Hoffman said he hopes the study will help experimentalists and data scientists who are designing AI models to consider how small variations in experimental design across labs can lead to problems with reproducibility and impact on long-term predictions for AI modes.

We see this work as a guide for the community as to how to think about designing experiments for inclusion in machine learning models,” Hoffman said.

This work was supported in part by the U.S. Department of Energy (DOE) Office of Science. Testing equipment was supplied in part by Co-ACCESS, part of the SUNCAT Center for Interface Science and Catalysis, a joint research center supported by SLAC National Accelerator Laboratory and Stanford University. The SLAC portion of the research took place at the Stanford Synchrotron Radiation Lightsource (SSRL), a DOE Office of Science user facility.

Tell Us What You Think

Do you have a review, update or anything you would like to add to this news story?

Leave your feedback
Your comment type
Submit

While we only use edited and approved content for Azthena answers, it may on occasions provide incorrect responses. Please confirm any data provided with the related suppliers or authors. We do not provide medical advice, if you search for medical information you must always consult a medical professional before acting on any information provided.

Your questions, but not your email details will be shared with OpenAI and retained for 30 days in accordance with their privacy principles.

Please do not ask questions that use sensitive or confidential information.

Read the full Terms & Conditions.