[2025-10-11 21:45:30,575][__main__][INFO] - Training for 50000 timesteps with NormalQNetwork and NormalReplayBuffer