{"id":1145,"date":"2025-06-16T04:33:26","date_gmt":"2025-06-16T04:33:26","guid":{"rendered":"https:\/\/synasc.ro\/2025\/?page_id=1145"},"modified":"2025-06-16T04:33:45","modified_gmt":"2025-06-16T04:33:45","slug":"synthetic-data-generation-with-llms","status":"publish","type":"page","link":"https:\/\/synasc.ro\/2025\/tutorials\/synthetic-data-generation-with-llms\/","title":{"rendered":"Synthetic Data Generation with LLMs"},"content":{"rendered":"\t\t<div data-elementor-type=\"wp-page\" data-elementor-id=\"1145\" class=\"elementor elementor-1145\" data-elementor-post-type=\"page\">\n\t\t\t\t<div class=\"elementor-element elementor-element-bef36a8 e-con-full e-flex e-con e-parent\" data-id=\"bef36a8\" data-element_type=\"container\">\n\t\t\t\t<div class=\"elementor-element elementor-element-9211027 elementor-widget elementor-widget-heading\" data-id=\"9211027\" data-element_type=\"widget\" data-widget_type=\"heading.default\">\n\t\t\t\t\t<h2 class=\"elementor-heading-title elementor-size-default\">Andreea Dutulescu, Stefan Ruseti, Mihai Dascalu<\/h2>\t\t\t\t<\/div>\n\t\t\t\t<div class=\"elementor-element elementor-element-9331b62 elementor-widget elementor-widget-text-editor\" data-id=\"9331b62\" data-element_type=\"widget\" data-widget_type=\"text-editor.default\">\n\t\t\t\t\t\t\t\t\t<div>National University of Science and Technology Politehnica Bucharest<\/div><div>\u00a0<\/div><div><b>Description<\/b>: This tutorial provides a technical overview of synthetic data generation using Large Language Models (LLMs), focusing on core methodologies and their integration. Synthetic data has become an essential tool for addressing key limitations in the availability, cost, and distributional coverage of manually annotated datasets. It enables scalable experimentation, facilitates data augmentation in low-resource settings, and supports iterative model refinement. The session begins with a discussion of generation methods and filtering strategies designed to enforce quality constraints. Next, the tutorial examines practical use cases. These include alignment tuning, where synthetic datasets are used to steer model behavior; inference-time augmentation, where generated exemplars support few-shot generalization or contextual adaptation; and self-improvement workflows, where models contribute to their iterative training through synthetic supervision.<\/div><div>\u00a0<\/div><div><em><strong>Short bios:<\/strong><\/em><\/div><ul><li><b>Andreea Dutulescu<\/b> is a PhD student at the National University of Science and Technology Politehnica Bucharest. Her research interests include natural language processing (NLP) in education, synthetic data generation, and related areas in applied machine learning. Andreea has gained experience in both academic and industrial settings. She has completed multiple internships at Google, where she worked on practical ML tasks and contributed to real-world applications. In her academic work, she has been involved in various research projects involving NLP.<\/li><li><b>\u0218tefan Ru\u0219e\u021bi<\/b> is an associate professor in the Department of Computer Science. His research activity spans multiple areas of Natural Language Processing (NLP), with most of his work focused on applying NLP techniques to develop tools for educational scenarios. He has experience in both national and international projects and has published over 80 papers, including 13 articles at top-tier conferences (EMNLP, COLING, ECIR, AIED) and 6 articles in Q1 journals (Computers &amp; Education, Computers in Human Behavior, International Journal of Artificial Intelligence in Education).<\/li><li><b>Mihai Dascalu<\/b>\u00a0is a professor of Computer Science at the National University of Science and Technology\u00a0 Politehnica\u00a0 Bucharest. He teaches object-oriented programming, algorithm design, and machine learning. He holds a dual Ph.D. in Computer Science and Educational Sciences and has authored over 300 research papers, including many in top-tier conferences and Q1 journals. Mihai has participated extensively in national and international projects, received prestigious awards such as a Fulbright scholarship, and is a corresponding member of the Academy of Romanian Scientists.<\/li><\/ul>\t\t\t\t\t\t\t\t<\/div>\n\t\t\t\t<\/div>\n\t\t\t\t<\/div>\n\t\t","protected":false},"excerpt":{"rendered":"<p>Andreea Dutulescu, Stefan Ruseti, Mihai Dascalu National University of Science and Technology Politehnica Bucharest\u00a0Description: This tutorial provides a technical overview of synthetic data generation using Large Language Models (LLMs), focusing on core methodologies and their integration. Synthetic data has become an essential tool for addressing key limitations in the availability, cost, and distributional coverage of [&hellip;]<\/p>\n","protected":false},"author":27,"featured_media":0,"parent":140,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-1145","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/synasc.ro\/2025\/wp-json\/wp\/v2\/pages\/1145","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/synasc.ro\/2025\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/synasc.ro\/2025\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/synasc.ro\/2025\/wp-json\/wp\/v2\/users\/27"}],"replies":[{"embeddable":true,"href":"https:\/\/synasc.ro\/2025\/wp-json\/wp\/v2\/comments?post=1145"}],"version-history":[{"count":4,"href":"https:\/\/synasc.ro\/2025\/wp-json\/wp\/v2\/pages\/1145\/revisions"}],"predecessor-version":[{"id":1149,"href":"https:\/\/synasc.ro\/2025\/wp-json\/wp\/v2\/pages\/1145\/revisions\/1149"}],"up":[{"embeddable":true,"href":"https:\/\/synasc.ro\/2025\/wp-json\/wp\/v2\/pages\/140"}],"wp:attachment":[{"href":"https:\/\/synasc.ro\/2025\/wp-json\/wp\/v2\/media?parent=1145"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}