Skip to content

About

In this environment, the agent navigates a 3x3 grid, trying to move from a starting position to a goal position. The agent receives rewards for reaching the goal and negative penalties for moving into an obstacle. The task is to learn an optimal policy (best sequence of actions) for reaching the goal with the highest cumulative reward.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

STANDALONE-RL

In this environment, the agent navigates a 3x3 grid, trying to move from a starting position to a goal position. The agent receives rewards for reaching the goal and negative penalties for moving into an obstacle. The task is to learn an optimal policy (best sequence of actions) for reaching the goal with the highest cumulative reward.

About

In this environment, the agent navigates a 3x3 grid, trying to move from a starting position to a goal position. The agent receives rewards for reaching the goal and negative penalties for moving into an obstacle. The task is to learn an optimal policy (best sequence of actions) for reaching the goal with the highest cumulative reward.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages