- Reworked the log_experiment flagging test to create sessions and runs directly through storage APIs. - Logged a baseline run, completed a second run, and then invoked log.execute using the baseline run ID in flag_runs. - Verified the baseline run was marked flagged with the expected reason via storage.listLoggedRuns output.