Repository for the analysis and preprocessing of datasets on the health status of electric vehicle batteries, specifically for battery capacity estimation.
This project is based on and builds upon the work presented by He et al. in their paper; more information in References.
For the sake of readers, a brief summary of the paper is also provided in this README.
- Project goal
- Dataset
- Repository structure
- Environments
- Installation from scratch
- Generating the five-fold files
The main objective of this project is to study the EVBattery dataset and establish a reproducible baseline which requires less datas for battery capacity estimation.
The authors formulate capacity estimation as a regression problem. Only charging snippets for which a capacity value is available are used in the capacity-estimation experiments.
The authors benchmark:
- Random Forest
- XGBoost
- MLP
- Gated CNN (GCNN)
- LSTM
The current implementation contains all of these branches in capacity_estimation/main.py. We will focus on Random Forest, XGBoost and LSTM later in this README.
For the thesis work, the authors' implementation is treated as a black box: the first objective is to reproduce and understand the original pipeline without making independent modifications.
The EVBattery dataset was collected from real-world electric vehicles from three manufacturers. The public release contains:
| Dataset | Vehicles | Anomaly vehicles | Charging snippets | Capacity labels |
|---|---|---|---|---|
battery_dataset1 |
217 | 31 | 629,121 | 349,741 |
battery_dataset2 |
198 | 1 | 472,829 | 203,207 |
battery_dataset3 |
49 | 16 | 176,327 | 32,974 |
The paper states that the capacity labels are real values in approximately the range 28.28–46.23 Ah.
The complete dataset contains more than 1.2 million charging snippets.
Each charging snippet contains 128 observations for eight time-series features:
- average cell voltage
- charging current
- SOC
- maximum cell voltage
- minimum cell voltage
- maximum cell temperature
- minimum cell temperature
- timestamp
The metadata associated with a snippet includes information such as vehicle number, mileage, charge-segment index, health label and capacity.
The raw public data is stored as pickle (.pkl) files.
To download the .zip with these datasets, see References.
The expected project layout is:
EVBattery/
│
├── original/
│ ├── battery_dataset1/
│ │ ├── data/
│ │ │ ├── *.pkl
│ │ │ └── ...
│ │ └── label/
│ │ └── label.csv
│ │
│ ├── battery_dataset2/
│ │ └── ...
│ │
│ └── battery_dataset3/
│ └── ...
│
├── processed/
│ ├── dataset1/
│ │ ├── data/
│ │ │ ├── *.dat
│ │ │ └── ...
│ │ └── metadata.pkl
│ │
│ ├── dataset2/
│ │ └── ...
│ │
│ └── dataset3/
│ └── ...
│
└── scripts/
├── authors_code/
│ └── battery_dataset_neurips23dataset_code/
│ │
│ ├── capacity_estimation/
│ │ ├── capacity_dataset.py
│ │ ├── main.py
│ │ └── ...
│ │
│ ├── five_fold_utils/
│ │ ├── all_car_dict.npz.npy
│ │ ├── ind_odd_dict.npz.npy
│ │ ├── ind_odd_dict1.npz.npy
│ │ ├── ind_odd_dict2.npz.npy
│ │ └── ind_odd_dict3.npz.npy
│ │
│ └── ...
│
└── my_code/
├── analyze_datasets.py
├── consolidate_datasets.py
├── inspect_pkl.py
└── ...
Important: the relative paths are significant for both authors' code and my code.
The processed/ dir will be automatically created by consolidate_datasets.py. You can check what each script in my_code/ does by opening the corresponding file: the first few lines provide an overview of its main purpose.
This project uses two separate Python environments:
The scripts developed for this project, located in my_code/, use:
- Python 3.11
- NumPy 2.4.6
The original code provided by the authors,located in authors_code/, requires a separate conda environment based on Python 3.6. The dependecies of their code are listed in the next section.
Important: The two environments should be kept separate because they require different Python and package versions.
The following procedure recreates the environment used for the capacity-estimation experiments.
Download and install Miniconda for Linux:
wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.shRun the installer:
bash Miniconda3-latest-Linux-x86_64.shFollow the installer instructions and restart the shell, or load Conda manually if requested by the installer.
Verify:
conda --versionThe authors' code is old and was released for Python 3.6.
Create the environment:
conda create -n evbattery python=3.6Activate it:
conda activate evbatteryVerify:
python --versionExpected:
Python 3.6.xx
Also verify which Python executable is being used:
which pythonExpected path is similar to:
/home/<user>/miniconda3/envs/evbattery/bin/python
As it's used just numpy, you can simply write with your env activated:
pip install numpyYou can install the exact versions used during with:
pip install -r requirements.txtImportant: Make sure that the evbattery conda environment is activated before installing the requirements. Otherwise, the packages may be installed into your system Python environment or another active environment.
I do not recommend doing the same with the requirements of the original authors' code, as they include also dependencies that are only needed for the battery health anomaly detection task, which are not covered in this repository.
Check CUDA availability:
python -c "import torch; print(torch.cuda.is_available())"If it prints False, you have to use the CPU-only setup of this project. Otherwise, you can uncomment the .cuda() calls in the main.py, in order to improve the performance.
COMING SOON
Move into the capacity-estimation directory:
cd ~/University/Tesi/EVBattery/scripts/authors_code/battery_dataset_neurips23dataset_code/capacity_estimationIf the repository was cloned somewhere else, use the corresponding absolute/relative path.
Make sure the Conda environment is active:
conda activate evbatteryThen run the default LSTM experiment:
python main.pyThe default arguments in the current main.py are:
fold_num = 0
batch_size = 64
model = LSTMNet
num_epochs = 10
Therefore:
python main.pyis equivalent to:
python main.py \
--fold_num 0 \
--batch_size 64 \
--model LSTMNet \
--num_epochs 10Selects the cross-validation fold.
Valid values for five-fold cross-validation are: 0, 1, 2, 3 or 4.
python main.py --fold_num 4Determines how many samples are processed at once during training. After processing one batch, the model uses the computed error to update its parameters.
python main.py --batch_size 32Selects the model between:
- LSTMNet
- MLP
- GatedCNN
- XGBoost
- RandomForest
- MEAN
python main.py --model XGBoostFor the LSTM/MLP/GatedCNN implementations, this is the number of training epochs.
For XGBoost, it is equivalent to the number of boosting rounds.
python main.py --num_epochs 50If present, the script loads previously serialized datasets from saved_dataset/ instead of rebuilding them.
python main.py --load_saved_datasetSome of the available models are the following:
The paper describes the LSTM architecture as an LSTM layer followed by two fully connected layers with ReLU activation, hidden dimension 32, Adam with learning rate 0.001, and 10 training epochs.
The implementation uses:
input_dim = 8
hidden_dim = 32
output_dim = 1
learning_rate = 0.001
loss = MSELoss
The current implementation uses:
objective = reg:squarederror
eta = 0.1
max_depth = 4
eval_metric = rmse
The number of boosting rounds is taken from --num_epochs and should be 50.
The current implementation uses:
n_estimators = 10
random_state = 0
n_jobs = 10
max_depth = 4
The main bottleneck is data loading and preprocessing.
The code iterates through many vehicle-associated .pkl files and filters them according to the capacity label.
On the CPU-only Ryzen 5 3500U machine used during this project, a single fold can take several minutes just to load/process the data before the actual training is complete.
Neural-network training is also significantly slower on CPU than on the NVIDIA GPUs used in the authors' experiments.
The capacity loader builds:
self.battery_dataset = []and appends every capacity-labeled snippet to the in-memory dataset. The dataset is therefore not streamed one file at a time during training. This can require substantial RAM.
On systems with limited RAM, the process may become very slow or the system may run out of memory while loading and preprocessing the data.
If you experience memory-related issues, monitor the system's RAM usage during execution and make sure that sufficient memory is available before running the full pipeline.
The project is based on the following paper:
Haowei He et al., EVBattery: A Large-Scale Electric Vehicle Dataset for Battery Health and Capacity Estimation, arXiv:2201.12358v3, 2023.
You can read the paper on arXiv and download the EVBattery Dataset on Figshare