[← projects]

Presented at CUCAI 2026

Benchmarking deep RL for off-grid hybrid microgrids

We tested six reinforcement learning algorithms on a simulated solar, battery, and diesel microgrid. The goal was simple: keep isolated communities powered while using as little fuel as possible.

150training runs
750ksteps per run
17,520decisions per year
2024held-out evaluation

The deployment problem

About 570 million people in sub-Saharan Africa lacked electricity access in 2022. For dispersed communities, extending the national grid can cost more than $3,000 per household. Local microgrids offer another path: solar for daytime demand, batteries for stored energy, and diesel when both fall short.

A microgrid is a small, local power system that can serve a community without relying on the national grid. In the system we studied, solar supplies the load first, the battery moves energy from sunny periods to darker ones, and a diesel generator provides backup when neither can meet demand.

These systems are usually controlled with fixed rules, such as starting diesel when the battery falls below a set charge level. The rules are simple, but they can waste fuel or become unreliable as weather and demand change. Model predictive control can make better-informed decisions, but it depends on forecasts that are difficult to maintain in data-scarce regions.

Reinforcement learning offers another option. A controller can learn in simulation, then adapt its decisions to the state of the system. Every 30 minutes, it must decide whether to store solar energy, discharge the battery, or run the generator. With no grid fallback, the question is simple:

Which deep RL methods can keep off-grid microgrids reliable across climates without wasting fuel?

The environment

We built a simulation for a small rural community. It models weather-driven solar output, battery efficiency and wear, diesel startup costs, minimum run times, and changing demand.

Solar50 m² array

18% efficiency, with temperature and soiling losses

Battery80 kWh

15-95% operating band, 20 kW dispatch

Diesel30 kW

$1.50/L, with startup and minimum-run limits

Figure 1

Physical topology of the simulated hybrid microgrid

WeatherNASA data
Demandload profile
Fuel200 L tank
Solar50 m² array
Battery80 kWh
Diesel30 kW
AC busbalances energy every 30 minutes
Served load
Curtailment
Unmet load
Weather and demand drive a shared AC bus. Solar is preferred, the battery moves energy across time, and diesel supplies firm backup when stored energy is not enough.

Solar power enters the AC bus first. The battery can store excess generation or discharge when solar falls short. Diesel is the final backup. Any remaining deficit becomes unmet load and receives the largest penalty in the system.

Figure 2

Six observations in, two continuous controls out

Observation · 6D
Battery SOC0 to 1
Fuel level0 to 1
Solar irradiance0 to 1.5 kW/m²
TemperatureT / 50, range -2 to 2
Time of day0 to 1
Load demand0 to 5 kWh
RL policy[256, 256] network
Action · 2D
Battery command
+1
discharge
0−1
charge
30% minimum load
battery-only
discharge
emergency
supply
solar
charging
pre-
conditioning
0 · off0.3
1 · full
Diesel throttle
The agent receives a compact state vector and returns battery dispatch and diesel throttle commands. Together, the two actions span four useful operating regions.

The algorithm sees six measurements and sets two continuous controls: battery power and diesel throttle. Those values produce charging, battery-only supply, pre-conditioning, or emergency generation.

How the reward works

The reward makes reliability the first constraint, then asks the agent to reduce operating cost and battery wear. Unmet energy costs 30 reward units per kWh, far more than the marginal fuel cost, so allowing a blackout cannot become an easy shortcut to a higher score.

Reward objective

Rt=−(puut+cfft+λcEFCt+psst+pggt)
pu = 30 $/kWhcf = 1.50 $/Lλc = 0.005ps = 0.50 $pg = 0.05 $/h
utunmet energy, kWh
ftdiesel used, L
EFCtbattery cycling
stgenerator start, 0 or 1
gtgenerator run time, h
Figure 3

Reward function decomposition

reward rt

Reliability

−30.0 × unmet kWhblackout penalty
+0.02 in SOC band20-90% buffer

Operating cost

−1.0 × fuel cost$1.50 per litre
−1.5 × curtailed kWhwasted solar
−0.5 per startcold-start wear
−0.05 × run timeoperating hours

Asset longevity

−0.005 × |battery energy|cycle wear
−0.0002 × |battery energy|degradation term
blackout penalty ≫ diesel cost ≫ battery wear
Reliability is intentionally weighted above operating cost and battery wear. This prevents an agent from appearing efficient by simply refusing to serve demand.

The equation gives the paper's formal objective. Figure 3 shows the fuller breakdown reported alongside it, including curtailed solar and a small bonus for keeping the battery between 20% and 90% charge.

A benchmark built to generalize

Six algorithms were trained across five locations, with five independent runs for each pairing. That produced 150 runs. Each algorithm trained on NASA POWER weather data from 2019 through 2023, while the full 2024 year stayed hidden until evaluation.

The sites span dry Sahel conditions, persistent equatorial cloud cover, coastal weather, and a highland monsoon. This makes it harder for an agent to succeed by memorizing one solar pattern.

Study map

Five climates across sub-Saharan Africa

DakarSenegal

Sahel / Atlantic coast

NiameyNiger

Semi-arid Sahel

KumasiGhana

Tropical rainforest

LibrevilleGabon

Equatorial humid

Addis AbabaEthiopia

Highland tropical

Tap, click, or focus a marker for the climate condition each location contributes to the benchmark.

What won

The off-policy methods were the clear winners. SAC, TQC, and DDPG kept unmet energy near zero, while every on-policy method failed in a distinct way. DDPG then separated itself by using 23% less fuel than SAC and TQC.

Their key advantage is a replay buffer. It stores past observations, actions, rewards, and resulting states so the algorithm can learn from them more than once. That gives rare blackouts and generator starts more weight during training, and helps connect earlier decisions with costs that appear much later. On-policy methods discard each batch after one update, making that connection harder to learn across a 17,520-step year.

7.5 kWhDDPG unmet energy / year
20,007 LDDPG diesel / year
23%less fuel than SAC and TQC
Table 1

Mean benchmark performance across 150 runs

Six algorithms, five locations, five training runs each. Held-out 2024 evaluation. Off-policy methods replay stored experience; on-policy methods learn from each rollout once. EFC = equivalent full battery cycles.

AlgorithmFamilyUnmet energy (kWh)Diesel (L)Battery EFC
DDPGOff-policy7.520,0072,335
SACOff-policy<126,0183,905
TQCOff-policy8.226,1103,407
A2COn-policy6,48918,50319
RPPOOn-policy3,35813,39425
PPOOn-policy11,28516525

Mean values from the released held-out evaluation data. Figure 4 shows the same values with variation across runs.

Figure 4

Benchmark performance across algorithms

A2CPPORPPOSACDDPGTQC

Episodic return

higher is better

reward

Unmet load

lower is better

kWh / year

Diesel use

lower is better

litres / year

Battery cycling

lower is better

EFC / year

Mean held-out 2024 performance across five locations and five independent training runs. Error bars show variation across runs.
01

DDPG adapted

It adjusted generator output to battery charge and local sunlight, keeping demand served with less fuel.

02

SAC and TQC stayed fixed

Both found a reliable diesel setting, but used about 26,000 litres in every climate.

03

The other methods were unreliable

PPO barely served demand, A2C varied widely, and recurrent PPO worked in only some training runs.

How learning unfolded

SAC, TQC, and DDPG found workable behavior early. PPO and A2C spent much more of training at low reward, while recurrent PPO improved only after a long, uneven recovery.

Figure 5

Training reward over 750,000 timesteps

Training reward

higher is better

Niamey · mean across five independent training runs

A2CPPORPPOSACDDPGTQC
Smoothed episodic return sampled from the released training runs. Off-policy methods settle early; the on-policy methods recover slowly or inconsistently.

The failure modes mattered

A single score can hide why an algorithm fails. PPO learned to do almost nothing: almost no diesel, almost no battery use, and about 11,500 kWh of unmet demand. It looked economical only because it stopped serving the load.

A2C avoided complete inaction but varied widely across training runs. Recurrent PPO showed a different problem: some runs learned a partly useful approach while others failed. One successful run would have made the method look much stronger than the full set did.

SAC and TQC were reliable but barely changed across climates. Their almost identical fuel use across five sites suggests one safe diesel setting rather than a response to local weather. DDPG was the only method that combined near-zero unmet load with meaningful fuel savings.

What this changed for us

The clearest result was that retaining experience mattered more than further reward tuning. Off-policy methods used replay buffers to learn from rare but costly events, including blackouts, generator starts, and poor battery decisions. On-policy methods discarded each rollout after updating, making it harder to connect early actions with costs that appeared much later. This helps explain why SAC, TQC, and DDPG learned reliable policies while PPO, A2C, and recurrent PPO remained unstable or failed.

For deployment, reliability must come first. A controller should consistently serve demand across climates and independent training runs before it is compared on fuel consumption and battery cycling. DDPG performed best by this standard: it maintained near-zero unmet demand, used less diesel than the other reliable methods, and adapted its behavior across climates. PPO’s low fuel use was not efficiency. It came from leaving demand unserved.

Beyond the results, this project was a valuable exercise in building and evaluating RL environments grounded in real-world data. Translating physical and operational data into realistic simulations is a broadly transferable skill. I hope to apply it to more problems in physics and business workflows; using these environments to benchmark agents, identify failure modes, and improve performance.

Presented at CUCAI 2026. The environment, training pipeline, and evaluation results are publicly available for reproduction and extension.

Explore the repository