JobbSafariLediga jobbMaster Thesis: LLM-based log processing and Anomaly Detection in Cloud Orchestartion Systems

Master Thesis: LLM-based log processing and Anomaly Detection in Cloud Orchestartion Systems
Ericsson ABSammanfattning
This thesis opportunity focuses on evaluating the effectiveness of Large Language Models (LLMs) in log processing and anomaly detection within large-scale cloud orchestration systems. The work will involve comparing LLM-based methods with traditional log parsing and machine learning techniques in a distributed environment, such as OpenStack or Kubernetes. The candidate will conduct research, design and implement processing pipelines, and analyze results to assess the applicability of LLMs for ITDet här erbjuder vi
Gain hands-on experience in cutting-edge technology and research.Contribute to the advancement of monitoring and diagnosis in distributed systems.
Ansök senast: Öppet tillsvidare
Publicerad: 2026-09-30
Beskrivning
Join our Team
About this opportunity
Large-scale cloud orchestration systems continuously generate vast volumes of log data from services responsible for managing compute, network, and storage resources. These logs contain valuable information about system health, service degradation, and emerging failures. Detecting and diagnosing such issues at an early stage is essential for maintaining reliable services and a positive user experience. In this thesis, you will investigate whether a Large Language Model (LLM) can perform log processing and anomaly detection as effectively as a traditional approach based on log parsing, such as LogDrain or Drain, combined with Machine Learning (ML) techniques. The work will focus on a complex, distributed cloud orchestration environment. Relevant examples include OpenStack and Kubernetes, although the specific target system will be selected in consultation with the thesis supervisor.
What you will do
The thesis will investigate how LLM-based approaches compare with conventional log-processing and ML pipelines in terms of: Identifying normal and healthy system behaviour. Detecting anomalous events and service degradation. Classifying and characterising fault scenarios. Identifying likely error causes or affected services. Accuracy, robustness, scalability, and computational cost.
As part of the thesis, you will:
Study relevant research on log parsing, anomaly detection, fault diagnosis, and LLM-based analysis of operational data.
Establish a representative baseline for healthy operation by collecting and analysing logs from the target system under normal conditions.
Design and implement one or more conventional log-processing pipelines using log parsing and ML techniques.
Design and implement an LLM-based approach for log processing and anomaly detection.
Define, construct, and inject representative fault scenarios into the target environment.
Evaluate how well the different approaches detect anomalies, identify fault types, and determine probable root causes.
Analyse the results and assess the practical applicability of LLMs for monitoring large-scale distributed systems.
The skills you bring
Ongoing Msc studies within Computer Science, Electrical Engineering, Physics or Maths
Knowledge of LLMs, AI, ML, Cloud
Expected outcome
The thesis will provide a systematic comparison of LLM-based and traditional approaches to log analysis. It will identify the strengths and limitations of each approach and provide recommendations for using LLMs in the monitoring and diagnosis of distributed cloud orchestration systems.
Suggested research question How effectively can an LLM process logs, detect anomalies, and support fault diagnosis in a distributed cloud orchestration system compared with a conventional pipeline based on log parsing and Machine Learning?
About this opportunity
Large-scale cloud orchestration systems continuously generate vast volumes of log data from services responsible for managing compute, network, and storage resources. These logs contain valuable information about system health, service degradation, and emerging failures. Detecting and diagnosing such issues at an early stage is essential for maintaining reliable services and a positive user experience. In this thesis, you will investigate whether a Large Language Model (LLM) can perform log processing and anomaly detection as effectively as a traditional approach based on log parsing, such as LogDrain or Drain, combined with Machine Learning (ML) techniques. The work will focus on a complex, distributed cloud orchestration environment. Relevant examples include OpenStack and Kubernetes, although the specific target system will be selected in consultation with the thesis supervisor.
What you will do
The thesis will investigate how LLM-based approaches compare with conventional log-processing and ML pipelines in terms of: Identifying normal and healthy system behaviour. Detecting anomalous events and service degradation. Classifying and characterising fault scenarios. Identifying likely error causes or affected services. Accuracy, robustness, scalability, and computational cost.
As part of the thesis, you will:
Study relevant research on log parsing, anomaly detection, fault diagnosis, and LLM-based analysis of operational data.
Establish a representative baseline for healthy operation by collecting and analysing logs from the target system under normal conditions.
Design and implement one or more conventional log-processing pipelines using log parsing and ML techniques.
Design and implement an LLM-based approach for log processing and anomaly detection.
Define, construct, and inject representative fault scenarios into the target environment.
Evaluate how well the different approaches detect anomalies, identify fault types, and determine probable root causes.
Analyse the results and assess the practical applicability of LLMs for monitoring large-scale distributed systems.
The skills you bring
Ongoing Msc studies within Computer Science, Electrical Engineering, Physics or Maths
Knowledge of LLMs, AI, ML, Cloud
Expected outcome
The thesis will provide a systematic comparison of LLM-based and traditional approaches to log analysis. It will identify the strengths and limitations of each approach and provide recommendations for using LLMs in the monitoring and diagnosis of distributed cloud orchestration systems.
Suggested research question How effectively can an LLM process logs, detect anomalies, and support fault diagnosis in a distributed cloud orchestration system compared with a conventional pipeline based on log parsing and Machine Learning?
Ansök till tjänsten
Master Thesis: LLM-based log processing and Anomaly Detection in Cloud Orchestartion Systems
OM FÖRETAGET

Ericsson AB










