{"title": "Bayesian models of human action understanding", "book": "Advances in Neural Information Processing Systems", "page_first": 99, "page_last": 106, "abstract": null, "full_text": "Bayesian models of human action understanding\n\nChris L. Baker, Joshua B. Tenenbaum & Rebecca R. Saxe\n\n{clbaker,jbt,saxe}@mit.edu\n\nDepartment of Brain and Cognitive Sciences\n\nMassachusetts Institute of Technology\n\nAbstract\n\nWe present a Bayesian framework for explaining how people reason\nabout and predict the actions of an intentional agent, based on observ-\ning its behavior. Action-understanding is cast as a problem of inverting\na probabilistic generative model, which assumes that agents tend to act\nrationally in order to achieve their goals given the constraints of their en-\nvironment. Working in a simple sprite-world domain, we show how this\nmodel can be used to infer the goal of an agent and predict how the agent\nwill act in novel situations or when environmental constraints change.\nThe model provides a qualitative account of several kinds of inferences\nthat preverbal infants have been shown to perform, and also \ufb01ts quantita-\ntive predictions that adult observers make in a new experiment.\n\n1 Introduction\n\nA woman is walking down the street. Suddenly, she turns 180 degrees and begins running\nin the opposite direction. Why? Did she suddenly realize she was going the wrong way,\nor change her mind about where she should be headed? Did she remember something\nimportant left behind? Did she see someone she is trying to avoid? These explanations for\nthe woman\u2019s behavior derive from taking the intentional stance: treating her as a rational\nagent whose behavior is governed by beliefs, desires or other mental states that refer to\nobjects, events, or states of the world [5].\n\nBoth adults and infants have been shown to make robust and rapid intentional inferences\nabout agents\u2019 behavior, even from highly impoverished stimuli. In \u201csprite-world\u201d displays,\nsimple shapes (e.g., circles) move in ways that convey a strong sense of agency to adults,\nand that lead to the formation of expectations consistent with goal-directed reasoning in in-\nfants [9, 8, 14]. The importance of the intentional stance in interpreting everyday situations,\ntogether with its robust engagement even in preverbal infants and with highly simpli\ufb01ed\nperceptual stimuli, suggest that it is a core capacity of human cognition.\n\nIn this paper we describe a computational framework for modeling intentional reasoning in\nadults and infants. Interpreting an agent\u2019s behavior via the intentional stance poses a highly\nunderconstrained inference problem: there are typically many con\ufb01gurations of beliefs and\ndesires consistent with any sequence of behavior. We de\ufb01ne a probabilistic generative\nmodel of an agent\u2019s behavior, in which behavior is dependent on hidden variables repre-\nsenting beliefs and desires. We then model intentional reasoning as a Bayesian inference\nabout these hidden variables given observed behavior sequences.\n\n\fIt is often said that \u201cvision is inverse graphics\u201d \u2013 the inversion of a causal physical process\nof scene formation. By analogy, our analysis of intentional reasoning might be called\n\u201cinverse planning\u201d, where the observer infers an agent\u2019s intentions, given observations of\nthe agent\u2019s behavior, by inverting a model of how intentions cause behavior. The intentional\nstance assumes that an agent\u2019s actions depend causally on mental states via the principle\nof rationality: rational agents tend to act to achieve their desires as optimally as possible,\ngiven their beliefs. To achieve their desired goals, agents must typically not only select\nsingle actions but must construct plans, or sequences of intended actions. The standards\nof \u201coptimal plan\u201d may vary with agent or circumstance: possibilities include achieving\ngoals \u201cas quickly as possible\u201d, \u201cas cheaply ...\u201d, \u201cas reliably ...\u201d, and so on. We assume a\nsoft, probabilistic version of the rationality principle, allowing that agents can often only\napproximate the optimal sequence of actions, and occasionally act in unexpected ways.\n\nThe paper is organized as follows. We \ufb01rst review several theoretical accounts of inten-\ntional reasoning from the cognitive science and arti\ufb01cial intelligence literatures, along\nwith some motivating empirical \ufb01ndings. We then present our computational framework,\ngrounding the discussion in a speci\ufb01c sprite-world domain. Lastly, we present results of\nour model on two sprite-world examples inspired by previous experiments in developmen-\ntal psychology, and results of the model on our own experiments.\n\n2 Empirical studies of intentional reasoning in infants and adults\n2.1 Inferring an invariant goal\nThe ability to predict how an agent\u2019s behavior will adapt when environmental circum-\nstances change, such as when an obstacle is inserted or removed, is a critical aspect of\nintentional reasoning. Gergely, Csibra and colleagues [8, 4] showed that preverbal infants\ncan infer an agent\u2019s goal that appears to be invariant across different circumstances, and\ncan predict the agent\u2019s future behavior by effectively assuming that it will act to achieve its\ngoal in an ef\ufb01cient way, subject to the constraints of its environment. Their experiments\nused a looking-time (violation-of-expectation) paradigm with sprite-world stimuli. Infant\nparticipants were assigned to one of two groups. In the \u201cobstacle\u201d condition, infants were\nhabituated to a sprite (a colored circle) moving (\u201cjumping\u201d) in a curved path over an ob-\nstacle to reach another object. The size of the obstacle varied across trials, but the sprite\nalways followed a near-shortest path over the obstacle to reach the other object. In the \u201cno\nobstacle\u201d group, infants were habituated to the sprite following the same curved \u201cjumping\u201d\ntrajectory to the other object, but without an obstacle blocking its path. Both groups were\nthen presented with the same test conditions, in which the obstacle was placed out of the\nsprite\u2019s way, and the sprite followed either the old, curved path or a new direct path to the\nother object. Infants from the \u201cobstacle\u201d group looked longer at the sprite following the\nunobstructed curved path, which (in the test condition) was now far from the most ef\ufb01cient\nroute to the other object. Infants in the \u201cno obstacle\u201d group looked equally at both test stim-\nuli. That is, infants in the \u201cobstacle\u201d condition appeared to interpret the sprite as moving in\na rational goal-directed fashion, with the other object as its goal. They expected the sprite\nto plan a path to the goal that was maximally ef\ufb01cient, subject to environmental constraints\nwhen present. Infants in the \u201cno obstacle\u201d group appeared more uncertain about whether\nthe sprite\u2019s movement was actually goal-directed or about what its goal was: was it simply\nto reach the other object, or something more complex, such as reaching the object via a\nparticular curved path?\n2.2 Inferring goals of varying complexity: rational means-ends analysis\nGergely et al. [6], expanding on work by Meltzoff [11], showed that infants can infer goals\nof varying complexity, again by interpreting agents\u2019 behaviors as rational responses to en-\nvironmental constraints. In two conditions, infants saw an adult demonstrate an unfamiliar\ncomplex action:\nIn the\n\u201chands occupied\u201d condition, the demonstrator pretended to be cold and wrapped a blanket\n\nilluminating a light-box by pressing its top with her forehead.\n\n\faround herself, so that she was incapable of using a more typical means (i.e., her hands)\nto achieve the same goal. In the \u201chands free\u201d condition the demonstrator had no such con-\nstraint. Most infants in the \u201chands free\u201d condition spontaneously performed the head-press\naction when shown the light-box one week later, but only a few infants in the \u201chands occu-\npied\u201d condition did so; the others illuminated the light-box simply by pressing it with their\nhands. Thus infants appear to assume that rational agents will take the most ef\ufb01cient path\nto their goal, and that if an agent appears to systematically employ an inef\ufb01cient means, it\nis likely because the agent has adopted a more complex goal that includes not only the end\nstate but also the means by which that end should be achieved.\n\n2.3 Inductive inference in intentional reasoning\nGergely and colleagues interpret their \ufb01ndings as if infants are reasoning about intentional\naction in an almost logical fashion, deducing the goal of an agent from its observed behav-\nior, the rationality principle, and other implicit premises. However, from a computational\npoint of view, it is surely oversimpli\ufb01ed to think that the intentional stance could be imple-\nmented in a deductive system. There are too many sources of uncertainty and the inference\nproblem is far too underconstrained for a logical approach to be successful. In contrast,\nour model posits that intentional reasoning is probabilistic. People\u2019s inferences about an\nagent\u2019s goal should be graded, re\ufb02ecting a tradeoff between the prior probability of a can-\ndidate goal and its likelihood in light of the agent\u2019s observed behavior. Inferences should\nbecome more con\ufb01dent as more of the agent\u2019s behavior is observed.\n\nTo test whether human intentional reasoning is consistent with a probabilistic account,\nit is necessary to collect data in greater quantities and with greater precision than infant\nstudies allow. Hence we designed our own sprite-world experimental paradigm, to collect\nricher quantitative judgments from adult observers. Many experiments are possible in this\nparadigm, but here we describe just one study of statistical effects on goal inference.\n\n(a)\n\nTraining paths\n\n(b)\n\nTest paths\n\n(c)\n\nd\nn\no\nc\n \n\ni\n\ng\nn\nn\na\nr\nT\n\ni\n\nl\n\ne\np\nm\ns\n\ni\n\nl\n\nx\ne\np\nm\no\nc\n\n1\n\n2\n\n3\n\n# stimuli\n\n4\n\n)\ns\np\nu\no\nr\ng\nh\n\n \n\nt\n\no\nb\n(\n\n1\n\n2\n\ng\nn\n\ni\nt\n\na\nr\n\n7\n\n5\n\n3\n\n1\n\nCond: simple\nCond: complex\n\n1\n\n2\n3\n# stimuli\n\n4\n\nFigure 1: (a) Training stimuli in complex and simple goal conditions. (b) Test stimuli 1 and 2. Test\nstimuli was the same for each group. (c) Mean of subjects\u2019 ratings with standard error bars (n=16).\nSixteen observers were told that they would be watching a series of animations of a mouse\nrunning in a simple maze (a box with a single internal wall). The displays were shown\nfrom an overhead perspective, with an animated schematic trace of the mouse\u2019s path as it\nran through the box. In each display, the mouse was placed in a different starting location\nand ran to recover a piece of cheese at a \ufb01xed, previously learned location. Observers were\ntold that the mouse had learned to follow a more-or-less direct path to the cheese, regardless\nof its starting location. Subjects saw two conditions in counterbalanced order. In one con-\ndition (\u201csimple goal\u201d), observers saw four displays consistent with this prior knowledge.\nIn another condition (\u201ccomplex goal\u201d), observers saw movements suggestive of a more\ncomplex, path-dependent goal for the mouse: it \ufb01rst ran directly to a particular location in\nthe middle of the box (the \u201cvia-point\u201d), and only then ran to the cheese. Fig. 1(a) shows\nthe mouse\u2019s four trajectories in each of these conditions. Note that the \ufb01rst trajectory was\nthe same in both conditions, while the next three were different. Also, all four trajecto-\nries in both conditions passed through the same hypothetical via-point in the middle of the\nbox, which was not marked in any conspicuous way. Hence both the simple goal (\u201cget to\n\n\fthe cheese\u201d) and complex goal (\u201cget to the cheese via point X\u201d) were logically possible\ninterpretations in both conditions.\n\nObservers\u2019 interpretations were assessed after viewing each of the four trajectories, by\nshowing them diagrams of two test paths (Fig. 1(b)) running from a novel starting location\nto the cheese. They were asked to rate the probability of the mouse taking one or the other\ntest path using a 1-7 scale: 1 = de\ufb01nitely path 1, 7 = de\ufb01nitely path 2, with intermediate\nvalues expressing intermediate degrees of con\ufb01dence. Observers in the simple-goal condi-\ntion always leaned towards path 1, the direct route that was consistent with the given prior\nknowledge. Observers in the complex-goal condition initially leaned just as much towards\npath 1, but after seeing additional trajectories they became increasingly con\ufb01dent that the\nmouse would follow path 2 (Fig. 1(c)). Importantly, the latter group increased its average\ncon\ufb01dence in path 2 with each subsequent trajectory viewed, consistent with the notion that\ngoal inference results from something like a Bayesian integration process: prior probability\nfavors the simple goal, but successive observations are more likely under the complex goal.\n\n3 Previous models of intentional reasoning\n\nThe above phenomena highlight two capacities than any model of intentional reasoning\nshould capture. First, representations of agents\u2019 mental states should include at least prim-\nitive planning capacities, with a constrained space of candidate goals and subgoals (or\nintended paths) that can refer to objects or locations in space, and the tendency to choose\naction sequences that achieve goals as ef\ufb01ciently as possible. Second, inferences about\nagents\u2019 goals should be probabilistic, and be sensitive to both prior knowledge about likely\ngoals as well as statistical evidence for more complex or less likely goals that better account\nfor observed actions.\n\nThese two components are clearly not suf\ufb01cient for a complete account of human inten-\ntional reasoning, but most previous accounts do not include even these capacities. Gergely,\nCsibra and colleagues [7] have proposed an informal (noncomputational) model in which\nagents are essentially treated as rational planners, but inferences about agents\u2019 goals are\npurely deductive, without a role for probabilistic expectations or gradations of con\ufb01dence.\n\nA more statistically sophisticated computational framework for inferring goals from behav-\nior has been proposed by [13], but this approach does not incorporate planning capacities.\nIn this framework, the observer learns to represent an agent\u2019s policies, conditional on the\nagent\u2019s goals. Within a static environment, this knowledge allows an observer to infer the\ngoal of an agent\u2019s actions, predict subsequent actions, and perform imitation, but it does\nnot support generalization to new environments where the agent\u2019s policy must adapt in re-\nsponse. Further, because generalization is not based on strong prior knowledge such as\nthe principle of rationality, many observations are needed for good performance. Likewise,\nprobabilistic approaches to plan recognition in AI (e.g., [3, 10]) typically represent plans\nin terms of policies (state-action pairs) that do not generalize when the structure of the\nenvironment changes in some unexpected way, and that require much data to learn from\nobservations of behavior.\n\nPerhaps closest to how people reason with the intentional stance are methods for inverse\nreinforcement learning (IRL) [12], or methods for learning an agent\u2019s utility function [2].\nBoth approaches assume a rational agent who maximizes expected utility, and attempt to\ninfer the agent\u2019s utility function from observations of its behavior. However, the utility\nfunctions that people attribute to intentional agents are typically much more structured\nand constrained than in conventional IRL. Goals are typically de\ufb01ned as relations towards\nobjects or other agents, and may include subgoals, preferred paths, or other elements. In the\nnext section we describe a Bayesian framework for modeling intentional reasoning that is\nsimilar in spirit to IRL, but more focused on the kinds of goal structures that are cognitively\nnatural to human adults and infants.\n\n\f4 The Bayesian framework\nWe propose to model intentional reasoning by combining the inferential power of statistical\napproaches to action understanding [12, 2, 13] with simple versions of the representational\nstructures that psychologists and philosophers [5, 7] have argued are essential in theory\nof mind. This section \ufb01rst presents our general approach, and then presents a speci\ufb01c\nmathematical model for the \u201cmouse\u201d sprite-world introduced above.\n\nMost generally, we assume a world that can be represented in terms of entities, attributes,\nand relations. Some attributes and relations are dynamic, indexed by a time dimension.\nSome entities are agents, who can perform actions at any time t with the potential to change\nthe world state at time t+1. We distinguish between environmental state, denoted W , and\nagent states, denoted S. For simplicity, we will assume that there is exactly one intentional\nagent in the world, and that the agent\u2019s actions can only affect its own state s \u2208 S. Let s0:T\nbe a sequence of T +1 agent states. Typically, observations of multiple state sequences of\nthe agent are available, and in general each may occur in a separate environment. Let s1:N\n0:T\nbe a set of N state sequences, and let w1:N be a set of N corresponding environments. Let\nAs be the set of actions available to the agent from state s, and let C(a) be the cost to the\nagent of action a \u2208 As. Let P (st+1|at, st, w) be the distribution over the agent\u2019s next state\nst+1, given the current state st, an action at \u2208 Ast, and the environmental state w.\nThe agent\u2019s actions are assumed to depend on mental states such as beliefs and desires. In\nour context, beliefs correspond to knowledge about the environmental state. Desires may\nbe simple or complex. A simple desire is an end goal: a world state or class of states that\nthe agent will act to bring about. There are many possibilities for more complex goals,\nsuch as achieving a certain end by means of a certain route, achieving a certain sequence\nof states in some order, and so on. We specify a particular goal space G of simple and\ncomplex goals for sprite-worlds in the next subsection. The agent draws goals g \u2208 G from\na prior distribution P (g|w1:N ), which constrains goals to be feasible in the environments\nw1:N from which observations of the agent\u2019s behavior are available.\nGiven the agent\u2019s goal g and an environment w, we can de\ufb01ne a value Vg,w(s) for each state\ns. The value function can be de\ufb01ned in various ways depending on the domain, task, and\nagent type. We specify a particular value function in the next subsection that re\ufb02ects the\ngoal structure of our sprite-world agent. The agent is assumed to choose actions according\nto a probabilistic policy, with a preference for actions with greater expected increases in\nvalue. Let Qg,w(s, a) = Ps\u2032 P (s\u2032|a, s, w)Vg,w(s\u2032) \u2212 C(a) be the expected value of the\nstate resulting from action a, minus the cost of the action. The agent\u2019s policy is\n\nP (at|st, g, w) \u221d exp(\u03b2Qg,w(st, at)).\n\n(1)\nThe parameter \u03b2 controls how likely the agent is to select the most valuable action. This\npolicy embodies a \u201csoft\u201d principle of rationality, which allows for inevitable sources of\nsuboptimal planning, or unexplained deviations from the direct path. A graphical model\nillustrating the relationship between the environmental state, and the agent\u2019s goals, actions,\nand states is shown in Fig. 2.\n\nThe observer\u2019s task is to infer g from the agent\u2019s behavior. We assume that state sequences\nare independent given the environment and the goal. The observer infers g from s1:N\n0:T via\nBayes\u2019 rule, conditional on w1:N :\n\nP (g|s1:N\n\n0:T , w1:N ) \u221d P (g|w1:N ) QN\n\n(2)\nWe assume that state transition probabilities and action probabilities are conditionally in-\ndependent given the agent\u2019s goal g, the agent\u2019s current state st, and the environment w.\nThe likelihood of a state sequence s0:T given a goal g and an environment w is computed\nby marginalizing over possible actions generating state transitions:\n\ni=1 P (si\n\n0:T |g, wi).\n\nP (s0:T |g, w) = QT \u22121\n\nt=0 Pat \u2208Ast\n\nP (st+1|at, st, w)P (at|st, g, w).\n\n(3)\n\n\fW\n\nG\n\n...\n\n...\n\nAt\n\nSt\n\nAt+1\n\nSt+1\n\n...\n...\n\nFigure 2: Two time-slice dynamic Bayes net representation\nof our model, where W is the environmental state, G is\nthe agent\u2019s goal, St is the agent\u2019s state at time t, and At\nis the agent\u2019s action at time t. Beliefs, desires, and actions\nintuitively map onto W , G and A, respectively.\n\n4.1 Modeling sprite-world inferences\nSeveral additional assumptions are necessary to apply the above framework to any speci\ufb01c\ndomain, such as the sprite-worlds discussed in \u00a72. The size of the grid, the location of\nobstacles, and likely goal points (such as the location of the cheese in our experimental\nstimuli) are represented by W , and assumed to be known to both the agent and the observer.\nThe agent\u2019s state space S consists of valid locations in the grid. All state sequences are\nassumed to be of the same length. The action space As consists of moves in all compass\ndirections {N, S, E, W, N E, N W, SE, SW }, except where blocked by an obstacle, and\naction costs are Euclidean. The agent can also choose to remain still with cost 1. We assume\nP (st+1|at, st, w) takes the agent to the desired adjacent grid point deterministically.\nThe set of possible goals G includes both simple and complex goals. Simple goals will just\nbe speci\ufb01c end states in S. While many kinds of complex goals are possible, we assume\nhere that a complex goal is just the combination of a desired end state with a desired means\nto achieving that end. In our sprite-worlds, we identify \u201cdesired means\u201d with a constraint\nthat the agent must pass through an additional speci\ufb01ed location enroute, such as the via-\npoint in the experiment from \u00a72.3. Because the number of complex goals de\ufb01ned in this\nway is much larger than the number of simple goals, the likelihood of each complex goal\nis small relative to the likelihood of individual simple goals. In addition, although path-\ndependent goals are possible, they should not be likely a priori. We thus set the prior\nP (g|w1:N ) to favor simple goals by a factor of \u03b3. For simplicity, we assume that the agent\ndraws just a single invariant goal g \u2208 G from P (g|w1:N ), and we assume that this prior\ndistribution is known to the observer. More generally, an agent\u2019s goals may vary across\ndifferent environments, and the prior P (g|w1:N ) may have to be learned.\nWe de\ufb01ne the value of a state Vg,w(s) as the expected total cost to the agent of achieving g\nwhile following the policy given in Eq. 1. We assume the desired end-state is absorbing and\ncost-free, which implies that the agent attempts the stochastic shortest path (with respect\nto its probabilistic policy) [1]. If g is a complex goal, Vg,w(s) is based on the stochastic\nshortest path through the speci\ufb01ed via-point. The agent\u2019s value function is computed using\nthe value iteration algorithm [1] with respect to the policy given in Eq. 1.\n\nFinally, to compare our model\u2019s predictions with behavioral data from human observers,\nwe must specify how to compute the probability of novel trajectories s\u2032\n0:T in a new envi-\nronment w\u2032, such as the test stimuli in Fig. 1, conditioned on an observed sequence s0:T in\nenvironment w. This is just an average over the predictions for each possible goal g:\n\nP (s\u2032\n\n0:T |s0:T , w, w\u2032) = Pg\u2208G P (s\u2032\n\n0:T |g, w\u2032)P (g|s0:T , w, w\u2032).\n\n(4)\n\n5 Sprite-world simulations\n5.1 Inferring an invariant goal\nAs a starting point for testing our model, we return to the experiments of Gergely et al. [8,\n4, 7], reviewed in \u00a72.1. Our input to the model, shown in Fig. 3(a,b), differs slightly from\nthe original stimuli used in [8], but the relevant details of interest are spared: goal-directed\naction in the presence of constraints. Our model predictions, shown in Fig. 3(c), capture\nthe qualitative results of these experiments, showing a large contrast between the straight\npath and the curved path in the condition with an obstacle, and a relatively small contrast\nin the condition with no obstacle. In the \u201cno obstacle\u201d condition, our model infers that the\nagent has a more complex goal, constrained by a via-point. This signi\ufb01cantly increases the\n\n\fprobability of the curved test path, to the point where the difference between the probability\nof observing curved and straight paths is negligible.\n\n(a)\n\nTraining paths\n\n(b)\n\nTest paths\n\n(c)\n\nTest: straight\nTest: curved\n\nd\nn\no\nc\n \ng\nn\nn\na\nr\nT\n\ni\n\ni\n\nt\ns\nb\no\n\nt\ns\nb\no\n\n \n\no\nn\n\n)\n)\nd\nn\no\nC\n\n|\nt\ns\ne\nT\n(\nP\n(\ng\no\n\u2212\n\nl\n\n20\n\n15\n\n10\n\n5\n\n0\n\n1\n\n1\n\n2\n\nCond: obst Cond: no obst\nFigure 3: Inferring an invariant goal. (a) Training input in obstacle and no obstacle conditions. (b)\nTest input is the same in each condition. (c) Model predictions: negative log likelihoods of test paths\n1 and 2 given data from training condition. In the obstacle condition, a large dissociation is seen\nbetween path 1 and path 2, with path 1 being much more likely. In the no obstacle condition, there is\nnot a large preference for either path 1 or path 2, qualitatively matching Gergely et al.\u2019s results [8].\n\n2\n\n5.2 Inferring goals of varying complexity: rational means-ends analysis\nOur next example is inspired by the studies of Gergely et al. [6] described in \u00a72.2.\nIn\nour sprite-world version of the experiment, we varied the amount of evidence for a simple\nversus a complex goal, by inputting the same three trajectories with and without an obstacle\npresent (Fig. 4(a)). In the \u201cobstacle\u201d condition, the trajectories were all approximately\nshortest paths to the goal, because the agent was forced to take indirect paths around the\nobstacle.\nIn the \u201cno obstacle\u201d condition, no such constraint was present to explain the\ncurved paths. Thus a more complex goal is inferred, with a path constrained to pass through\na via-point. Given a choice of test paths, shown in Fig. 4(b), the model shows a double-\ndissociation between the probability of the direct path and the curved path through the\nputative via-point, given each training condition (Fig. 4(c)), similar to the results in [6].\n\n(a)\n\nTraining paths\n\n(b)\n\nTest paths\n\nd\nn\no\nc\n \ng\nn\nn\na\nr\nT\n\ni\n\ni\n\nt\ns\nb\no\n\nt\ns\nb\no\n\n \n\no\nn\n\nTest: straight\nTest: curved\n\n20\n\n15\n\n10\n\n5\n\n(c)\n\n)\n)\nd\nn\no\nC\n\n|\nt\ns\ne\nT\n(\nP\n(\ng\no\n\u2212\n\nl\n\n1\n\n2\n\n3\n\nCond: obst Cond: no obst\nFigure 4: Inferring goals of varying complexity. (a) Training input in obstacle and no obstacle con-\nditions.\n(c) Model predictions: a double dissociation between\nprobability of test paths 1 and 2 in the two conditions. This re\ufb02ects a preference for the straight path\nin the \ufb01rst condition, where there is an obstacle to explain the agent\u2019s de\ufb02ections in the training input,\nand a preference for the curved path in the second condition, where a complex goal is inferred.\n\n(b) Test input in each condition.\n\n1\n\n2\n\n0\n\n5.3 Inductive inference in intentional reasoning\nLastly, we present the results of our model on our own behavioral experiment, \ufb01rst de-\nscribed in \u00a72.3 and shown in Fig. 1. These data demonstrated the statistical nature of peo-\nple\u2019s intentional inferences. Fig. 5 compares people\u2019s judgments of the probability that the\nagent takes a particular test path with our model\u2019s predictions. To place model predictions\nand human judgments on a comparable scale, we \ufb01t a sigmoidal psychometric transforma-\ntion to the computed log posterior odds for the curved test path versus the straight path. The\nBayesian model captures the graded shift in people\u2019s expectations in the \u201ccomplex goal\u201d\ncondition, as evidence accumulates that the agent always seeks to pass through an arbitrary\nvia-point enroute to the end state.\n\n\fCond: simple\nCond: complex\n\n7\n\n5\n\n3\n\ng\nn\ni\nt\na\nr\n\n1\n\n1\n\nFigure 5: Experimental results: model \ufb01t for behavioral data.\nMean ratings are plotted as hollow circles. Error bars give standard\nerror. The log posterior odds from the model were \ufb01t to subjects\u2019\nratings using a scaled sigmoid function with range (1, 7). The sig-\nmoid function includes bias and gain parameters, which were \ufb01t to\nthe human data by minimizing the sum-squared error between the\nmodel predictions and mean subject ratings.\n\n4\n\n2\n3\n# stimuli\n6 Conclusion\nWe presented a Bayesian framework to explain several core aspects of intentional reason-\ning: inferring the goal of an agent based on observations of its behavior, and predicting\nhow the agent will act when constraints or initial conditions for action change. Our model\ncaptured basic qualitative inferences that even preverbal infants have been shown to per-\nform, as well as more subtle quantitative inferences that adult observers made in a novel\nexperiment. Two future challenges for our computational framework are: representing and\nlearning multiple agent types (e.g. rational, irrational, random, etc.), and representing and\nlearning hierarchically structured goal spaces that vary across environments, situations and\neven domains. These extensions will allow us to further test the power of our computational\nframework, and will support its application to the wide range of intentional inferences that\npeople constantly make in their everyday lives.\nAcknowledgments: We thank Whitman Richards, Konrad K\u00a8ording, Kobi Gal, Vikash Mans-\ninghka, Charles Kemp, and Pat Shafto for helpful comments and discussions.\nReferences\n[1] D. P. Bertsekas. Dynamic Programming and Optimal Control. Athena Scienti\ufb01c, Belmont,\n\nMA, 2nd edition, 2001.\n\n[2] U. Chajewska, D. Koller, and D. Ormoneit. Learning an agent\u2019s utility function by observing\n\nbehavior. In Proc. of the 18th Intl. Conf. on Machine Learning (ICML), pages 35\u201342, 2001.\n\n[3] E. Charniak and R. Goldman. A probabilistic model of plan recognition. In Proc. AAAI, 1991.\n[4] G. Csibra, G. Gergely, S. Bir\u00b4o, O. Ko\u00b4os, and M. Brockbank. Goal attribution without agency\n\ncues: the perception of \u2018pure reason\u2019 in infancy. Cognition, 72:237\u2013267, 1999.\n\n[5] D. C. Dennett. The Intentional Stance. Cambridge, MA: MIT Press, 1987.\n[6] G. Gergely, H. Bekkering, and I. Kir\u00b4aly. Rational imitation in preverbal infants. Nature,\n\n415:755, 2002.\n\n[7] G. Gergely and G. Csibra. Teleological reasoning in infancy: the na\u00a8\u0131ve theory of rational action.\n\nTrends in Cognitive Sciences, 7(7):287\u2013292, 2003.\n\n[8] G. Gergely, Z. N\u00b4adasdy, G. Csibra, and S. Bir\u00b4o. Taking the intentional stance at 12 months of\n\nage. Cognition, 56:165\u2013193, 1995.\n\n[9] F. Heider and M. A. Simmel. An experimental study of apparent behavior. American Journal\n\nof Psychology, 57:243\u2013249, 1944.\n\n[10] L. Liao, D. Fox, and H. Kautz. Learning and inferring transportation routines. In Proc. AAAI,\n\npages 348\u2013353, 2004.\n\n[11] A. N. Meltzoff. Infant imitation after a 1-week delay: Long-term memory for novel acts and\n\nmultiple stimuli. Developmental Psychology, 24:470\u2013476, 1988.\n\n[12] A. Y. Ng and S. Russell. Algorithms for inverse reinforcement learning. In Proc. of the 17th\n\nIntl. Conf. on Machine Learning (ICML), pages 663\u2013670, 2000.\n\n[13] R. P. N. Rao, A. P. Shon, and A. N. Meltzoff. A Bayesian model of imitation in infants and\n\nrobots. In Imitation and Social Learning in Robots, Humans, and Animals. (in press).\n\n[14] B. J. Scholl and P. D. Tremoulet. Perceptual causality and animacy. Trends in Cognitive Sci-\n\nences, 4(8):299\u2013309, 2000.\n\n\f", "award": [], "sourceid": 2815, "authors": [{"given_name": "Chris", "family_name": "Baker", "institution": null}, {"given_name": "Rebecca", "family_name": "Saxe", "institution": null}, {"given_name": "Joshua", "family_name": "Tenenbaum", "institution": null}]}