Shaul Eliahou-Niv
doi.org/10.36647/CIML/07.01.A003
Abstract : In this work, we evaluate a relatively small, general‑purpose large language model (LLM) on a specialized and complex Systems Engineering task. Specifically, we employ a Systems Engineering benchmark to accurately measure the model’s performance. To improve accuracy and reduce hallucinations, we apply a Retrieval‑Augmented Generation (RAG) framework based on a concise corpus composed of Systems Engineering–related textbooks. In addition, we investigate the impact of the RAG corpus on the accuracy of the results. We find that when RAG is combined with a subject‑specific benchmark, the framework can effectively act as a judge ranking texts by their relevance to the subject and, when relevant, assessing their completeness relative to the full scope of the benchmark.
Keyword : Artificial intelligence, large language model, retrieval augmented generation, systems engineering.