Abstract
Mark K. Ho
Abstract
Authors
Institutions
Provenance
crossref
Confidence 100%
pubmed
Confidence 98%
europepmc
Confidence 96%
openalex
Confidence 95%
datacite
Confidence 0%
No local reference links have been materialized yet.
No local citing links have been materialized yet.
10.52202/075280-2192
10.52202/075280-2192
10.24963/ijcai.2022/730
10.24963/ijcai.2022/730
10.1201/9781315140223
10.1201/9781315140223 · 2021
Unresolved referenced work
1990
Utility theory without the completeness axiom
10.2307/1909888 · 1962
Optimizing return distributions with distributional dynamic programming
2025
Unresolved referenced work
Kept as external metadata until matched
Constructing future behavior in the hippocampal formation through composition and replay
10.1038/s41593-025-01908-3 · 2025
Unresolved referenced work
2023
Unresolved referenced work
Kept as external metadata until matched
10.7551/mitpress/14207.001.0001
10.7551/mitpress/14207.001.0001 · 2023
A Markovian decision process
1957
Parsing reward
10.1016/s0166-2236(03)00233-9 · 2003
Discrete dynamic programming
10.1214/aoms/1177704593 · 1962
Unresolved referenced work
Kept as external metadata until matched
Unresolved referenced work
1987
Unresolved referenced work
2015
10.1002/9781118920497.ch1
10.1002/9781118920497.ch1
10.1609/aaai.v32i1.11791
10.1609/aaai.v32i1.11791
Unresolved referenced work
Kept as external metadata until matched
Goals as reward‐producing programs
10.1038/s42256-025-00981-4 · 2025
Reinforcement learning: The good, the bad and the ugly
10.1016/j.conb.2008.08.003 · 2008
Unresolved referenced work
Kept as external metadata until matched
Unresolved referenced work
1989
Actions and habits: The development of behavioural autonomy
10.1098/rstb.1985.0010 · 1985
10.7551/mitpress/2872.003.0015
10.7551/mitpress/2872.003.0015 · 2000
Simple agent, complex environment: Efficient reinforcement learning with agent states
2022
Connectionism and cognitive architecture: A critical analysis
10.1016/0010-0277(88)90031-5 · 1988
Unresolved referenced work
Kept as external metadata until matched
Unresolved referenced work
Kept as external metadata until matched
How toddlers begin to learn verbs
10.1016/j.tics.2008.07.003 · 2008
Doing more with less: Meta‐reasoning and meta‐learning in humans and machines
10.1016/j.cobeha.2019.01.005 · 2019
10.7551/mitpress/8765.001.0001
10.7551/mitpress/8765.001.0001 · 2024
The value equivalence principle for model‐based reinforcement learning
2020
10.65109/qtwt7869
10.65109/qtwt7869
People construct simplified mental representations to plan
10.1038/s41586-022-04743-9 · 2022
Rational simplification and rigidity in human planning
10.1177/09567976231200547 · 2023
People teach with rewards and punishments as communication, not reinforcements
10.1037/xge0000569 · 2019
Risk‐sensitive Markov decision processes
10.1287/mnsc.18.7.356 · 1972
Unresolved referenced work
Kept as external metadata until matched
Fuzzy logic
10.1109/2.53 · doi-reference
Psychology of habit
10.1146/annurev-psych-122414-033417 · doi-reference
Rational choice and the structure of the environment
10.1037/h0042769 · doi-reference
Toward a rational and mechanistic account of mental effort
10.1146/annurev-neuro-072116-031526 · doi-reference
How intractability spans the cognitive and evolutionary levels of explanation
10.1111/tops.12506 · doi-reference
The logical paradoxes and the law of excluded middle
10.2307/2218742 · doi-reference
Does the chimpanzee have a theory of mind?
10.1017/s0140525x00076512 · doi-reference
An integrative theory of prefrontal cortex function
10.1146/annurev.neuro.24.1.167 · doi-reference
Risk‐sensitive reinforcement learning
10.1023/a:1017940631555 · doi-reference
Planning in the brain
10.1016/j.neuron.2021.12.018 · doi-reference
Average reward reinforcement learning: Foundations, algorithms, and empirical results
10.1023/a:1018064306595 · doi-reference
Self‐regulation through goal setting
10.1016/0749-5978(91)90021-k · doi-reference
Building machines that learn and think like people
10.1017/s0140525x16001837 · doi-reference
Compression and communication in the cultural evolution of linguistic structure
10.1016/j.cognition.2015.03.016 · doi-reference
Planning and acting in partially observable stochastic domains
10.1016/s0004-3702(98)00023-x · doi-reference
The naïve utility calculus: Computational principles underlying commonsense psychology
10.1016/j.tics.2016.05.011 · doi-reference
Reward machines: Exploiting reward function structure in reinforcement learning
10.1613/jair.1.12440 · doi-reference
Risk‐sensitive Markov decision processes
10.1287/mnsc.18.7.356 · doi-reference
People teach with rewards and punishments as communication, not reinforcements
10.1037/xge0000569 · doi-reference
Rational simplification and rigidity in human planning
10.1177/09567976231200547 · doi-reference
People construct simplified mental representations to plan
10.1038/s41586-022-04743-9 · doi-reference
10.65109/qtwt7869
10.65109/qtwt7869 · doi-reference
10.7551/mitpress/8765.001.0001
10.7551/mitpress/8765.001.0001 · doi-reference
Doing more with less: Meta‐reasoning and meta‐learning in humans and machines
10.1016/j.cobeha.2019.01.005 · doi-reference
How toddlers begin to learn verbs
10.1016/j.tics.2008.07.003 · doi-reference
Connectionism and cognitive architecture: A critical analysis
10.1016/0010-0277(88)90031-5 · doi-reference
10.7551/mitpress/2872.003.0015
10.7551/mitpress/2872.003.0015 · doi-reference
Actions and habits: The development of behavioural autonomy
10.1098/rstb.1985.0010 · doi-reference
Reinforcement learning: The good, the bad and the ugly
10.1016/j.conb.2008.08.003 · doi-reference
Goals as reward‐producing programs
10.1038/s42256-025-00981-4 · doi-reference
10.1609/aaai.v32i1.11791
10.1609/aaai.v32i1.11791 · doi-reference
10.1002/9781118920497.ch1
10.1002/9781118920497.ch1 · doi-reference
Discrete dynamic programming
10.1214/aoms/1177704593 · doi-reference
Parsing reward
10.1016/s0166-2236(03)00233-9 · doi-reference
10.7551/mitpress/14207.001.0001
10.7551/mitpress/14207.001.0001 · doi-reference
Constructing future behavior in the hippocampal formation through composition and replay
10.1038/s41593-025-01908-3 · doi-reference
Utility theory without the completeness axiom
10.2307/1909888 · doi-reference
10.1201/9781315140223
10.1201/9781315140223 · doi-reference
10.24963/ijcai.2022/730
10.24963/ijcai.2022/730 · doi-reference
10.52202/075280-2192
10.52202/075280-2192 · doi-reference