Presented at CUCAI 2026
Benchmarking deep RL for off-grid hybrid microgrids
We tested six reinforcement learning algorithms on a simulated solar, battery, and diesel microgrid. The goal was simple: keep isolated communities powered while using as little fuel as possible.
The deployment problem
About 570 million people in sub-Saharan Africa lacked electricity access in 2022. For dispersed communities, extending the national grid can cost more than $3,000 per household. Local microgrids offer another path: solar for daytime demand, batteries for stored energy, and diesel when both fall short.
A microgrid is a small, local power system that can serve a community without relying on the national grid. In the system we studied, solar supplies the load first, the battery moves energy from sunny periods to darker ones, and a diesel generator provides backup when neither can meet demand.
These systems are usually controlled with fixed rules, such as starting diesel when the battery falls below a set charge level. The rules are simple, but they can waste fuel or become unreliable as weather and demand change. Model predictive control can make better-informed decisions, but it depends on forecasts that are difficult to maintain in data-scarce regions.
Reinforcement learning offers another option. A controller can learn in simulation, then adapt its decisions to the state of the system. Every 30 minutes, it must decide whether to store solar energy, discharge the battery, or run the generator. With no grid fallback, the question is simple:
Which deep RL methods can keep off-grid microgrids reliable across climates without wasting fuel?
The environment
We built a simulation for a small rural community. It models weather-driven solar output, battery efficiency and wear, diesel startup costs, minimum run times, and changing demand.
18% efficiency, with temperature and soiling losses
15-95% operating band, 20 kW dispatch
$1.50/L, with startup and minimum-run limits
Physical topology of the simulated hybrid microgrid
Solar power enters the AC bus first. The battery can store excess generation or discharge when solar falls short. Diesel is the final backup. Any remaining deficit becomes unmet load and receives the largest penalty in the system.
Six observations in, two continuous controls out
discharge0−1
charge
discharge
supply
charging
conditioning
The algorithm sees six measurements and sets two continuous controls: battery power and diesel throttle. Those values produce charging, battery-only supply, pre-conditioning, or emergency generation.
How the reward works
The reward makes reliability the first constraint, then asks the agent to reduce operating cost and battery wear. Unmet energy costs 30 reward units per kWh, far more than the marginal fuel cost, so allowing a blackout cannot become an easy shortcut to a higher score.
Reward objective
Reward function decomposition
Reliability
Operating cost
Asset longevity
The equation gives the paper's formal objective. Figure 3 shows the fuller breakdown reported alongside it, including curtailed solar and a small bonus for keeping the battery between 20% and 90% charge.
A benchmark built to generalize
Six algorithms were trained across five locations, with five independent runs for each pairing. That produced 150 runs. Each algorithm trained on NASA POWER weather data from 2019 through 2023, while the full 2024 year stayed hidden until evaluation.
The sites span dry Sahel conditions, persistent equatorial cloud cover, coastal weather, and a highland monsoon. This makes it harder for an agent to succeed by memorizing one solar pattern.
Five climates across sub-Saharan Africa
Sahel / Atlantic coast
Semi-arid Sahel
Tropical rainforest
Equatorial humid
Highland tropical
What won
The off-policy methods were the clear winners. SAC, TQC, and DDPG kept unmet energy near zero, while every on-policy method failed in a distinct way. DDPG then separated itself by using 23% less fuel than SAC and TQC.
Their key advantage is a replay buffer. It stores past observations, actions, rewards, and resulting states so the algorithm can learn from them more than once. That gives rare blackouts and generator starts more weight during training, and helps connect earlier decisions with costs that appear much later. On-policy methods discard each batch after one update, making that connection harder to learn across a 17,520-step year.
Mean benchmark performance across 150 runs
Six algorithms, five locations, five training runs each. Held-out 2024 evaluation. Off-policy methods replay stored experience; on-policy methods learn from each rollout once. EFC = equivalent full battery cycles.
| Algorithm | Family | Unmet energy (kWh) | Diesel (L) | Battery EFC |
|---|---|---|---|---|
| DDPG | Off-policy | 7.5 | 20,007 | 2,335 |
| SAC | Off-policy | <1 | 26,018 | 3,905 |
| TQC | Off-policy | 8.2 | 26,110 | 3,407 |
| A2C | On-policy | 6,489 | 18,503 | 19 |
| RPPO | On-policy | 3,358 | 13,394 | 25 |
| PPO | On-policy | 11,285 | 165 | 25 |
Mean values from the released held-out evaluation data. Figure 4 shows the same values with variation across runs.
Figure 4
Benchmark performance across algorithms
Episodic return
higher is betterreward
Unmet load
lower is betterkWh / year
Diesel use
lower is betterlitres / year
Battery cycling
lower is betterEFC / year
DDPG adapted
It adjusted generator output to battery charge and local sunlight, keeping demand served with less fuel.
SAC and TQC stayed fixed
Both found a reliable diesel setting, but used about 26,000 litres in every climate.
The other methods were unreliable
PPO barely served demand, A2C varied widely, and recurrent PPO worked in only some training runs.
How learning unfolded
SAC, TQC, and DDPG found workable behavior early. PPO and A2C spent much more of training at low reward, while recurrent PPO improved only after a long, uneven recovery.
Figure 5
Training reward over 750,000 timesteps
Training reward
higher is betterNiamey · mean across five independent training runs
The failure modes mattered
A single score can hide why an algorithm fails. PPO learned to do almost nothing: almost no diesel, almost no battery use, and about 11,500 kWh of unmet demand. It looked economical only because it stopped serving the load.
A2C avoided complete inaction but varied widely across training runs. Recurrent PPO showed a different problem: some runs learned a partly useful approach while others failed. One successful run would have made the method look much stronger than the full set did.
SAC and TQC were reliable but barely changed across climates. Their almost identical fuel use across five sites suggests one safe diesel setting rather than a response to local weather. DDPG was the only method that combined near-zero unmet load with meaningful fuel savings.
What this changed for us
The clearest result was that retaining experience mattered more than further reward tuning. Off-policy methods used replay buffers to learn from rare but costly events, including blackouts, generator starts, and poor battery decisions. On-policy methods discarded each rollout after updating, making it harder to connect early actions with costs that appeared much later. This helps explain why SAC, TQC, and DDPG learned reliable policies while PPO, A2C, and recurrent PPO remained unstable or failed.
For deployment, reliability must come first. A controller should consistently serve demand across climates and independent training runs before it is compared on fuel consumption and battery cycling. DDPG performed best by this standard: it maintained near-zero unmet demand, used less diesel than the other reliable methods, and adapted its behavior across climates. PPO’s low fuel use was not efficiency. It came from leaving demand unserved.
Beyond the results, this project was a valuable exercise in building and evaluating RL environments grounded in real-world data. Translating physical and operational data into realistic simulations is a broadly transferable skill. I hope to apply it to more problems in physics and business workflows; using these environments to benchmark agents, identify failure modes, and improve performance.
Presented at CUCAI 2026. The environment, training pipeline, and evaluation results are publicly available for reproduction and extension.
Explore the repository