Repository navigation
A question regarding ReinforceLoss and ProbabilisticActor behaviour #3378
|
Hi, thank you for maintaining this project. I might be misunderstanding something, so I wanted to ask for clarification. TL;DR
SetupI was trying to setup a simple experiment using a combination of ProbabilisticActor and ReinforceLoss. For simplicity, I used data obtained from a single EnvBase.rollout(). While reading the docs and code, I noticed the following points:
Expected behavior (my understanding)Based on the above, I expected the following:
What I observedIn practice, I have observed the opposite behaviour:
This was surprising to me, as it seems to imply that Looking further into the QuestionIs this behavior intended, or am I missing some assumption or configuration detail? I also noticed that other losses (e.g. Code / referencesI have also conducted a small test which demonstrates the behaviour I am describing: Output of the script: |
Replies: 1 comment
|
@ArturGoldman I just reproduced this scenario and it seems that there is bug where the rollouts get overwritten with new samples. I can draft a PR fixing this issue for the REINFORCE Loss. I'll also update the docstring for the ProbabilisticActor behavior in this pr. |
@ArturGoldman I just reproduced this scenario and it seems that there is bug where the rollouts get overwritten with new samples. I can draft a PR fixing this issue for the REINFORCE Loss. I'll also update the docstring for the ProbabilisticActor behavior in this pr.