{"title": "Learning direction in global motion: two classes of psychophysically-motivated models", "book": "Advances in Neural Information Processing Systems", "page_first": 917, "page_last": 924, "abstract": null, "full_text": "Learning direction in global motion:  two \n\nclasses of psychophysically-motivated \n\nmodels \n\nV.  Sundareswaran \n\nLucia M. Vaina* \n\nIntelligent Systems Laboratory, College of Engineering, \n\nBoston University \n\n44 Cummington Street, Boston,  MA 02215 \n\nAbstract \n\nPerceptual learning is  defined  as  fast  improvement  in  performance and \nretention  of the  learned  ability  over  a  period  of time.  In a  set  of psy(cid:173)\nchophysical experiments  we  demonstrated  that  perceptual learning oc(cid:173)\ncurs for the discrimination of direction in stochastic motion stimuli.  Here \nwe  model  this  learning  using  two  approaches:  a clustering  model  that \nlearns  to  accommodate  the motion  noise,  and an averaging  model  that \nlearns to  ignore the noise.  Simulations of the models show  performance \nsimilar to the psychophysical results. \n\n1 \n\nIntroduction \n\nGlobal  motion  perception is  critical to  many visual tasks:  to perceive self-motion, \nto  identify objects in  motion,  to determine  the structure of the environment,  and \nto make judgements for  safe navigation.  In the presence of noise,  as in random dot \nkinematograms,  efficient  extraction  of global  motion  involves  considerable  spatial \nintegration.  Newsome and Colleagues  (1989)  showed that neurons in the macaque \nmiddle temporal area (MT) are motion direction-selective, and perform global inte(cid:173)\ngration of motion in their  large receptive fields.  Psychophysical studies in humans \nhave characterized the limits of spatial and temporal integration in motion  (Wata(cid:173)\nmaniuk et. aI,  1984)  and the nature of the underlying motion computations (Vaina \net.  al  1990). \n\n\u00b7Please address all correspondence to Lucia Vaina \n\n\f918 \n\nV.  Sundareswaran,  Lucia M.  Vaina \n\nSince  the psychophysical  and neural substrate of global  motion  are fairly  well  un(cid:173)\nderstood,  we  were  interested  to  see  whether  the  perception  of  direction  in  such \nglobal  motion  stimuli  can  improve  with  practice.  Studies  specifically  addressing \nthis question for other early perceptual tasks have shown that improvements of per(cid:173)\nformance  obtained  in  the  first  experimental session  are  preserved  in  a  subsequent \nsession and retained over weeks.  This is considered as perceptual learning (Gibson, \n1953).  Psychophysical studies of perceptual learning show that the beneficial effects \nof practice  are  lost  if some  stimulus  parameters  are changed  significantly,  such  as \norientation, spatial frequency or location in the visual field.  Based on the time scale \nnecessary  for  the improvement  to occur,  two  major  learning paradigms  have  been \nused in  perceptual learning :  slow,  progressive learning (several thousand trials are \nrequired to  reach  stable performance)  and fast  !earning  (improvement  occurs  and \nstabilizes in the first  100-200 trials). \nThe idea of fast  learning  and the  nature of its  limits  is  attractive from  a  compu(cid:173)\ntational point of view  because it  encourages the exploration of practice-dependent \nplasticity found  in the adult early visual system  (Fregnac et.  al.,  1988,  Gilbert and \nWiesel  1992).  A  recent  line  of research  in  biologically  motivated  learning  models \noriginated by Poggio  (Poggio,  1990)  takes perceptual learning as evidence that\"the \nbrain  may  be  able  to  synthesize-possibly  in  the  cortex-appropriate  task-specific \nmodules  that  receive  input  from  retinotopic  cells  and learn to solve  the task after \na short training phase in which they are exposed to examples of the task\".  Poggio \nand colleagues (1992)  have illustrated this approach in learning vernier hyperacuity. \nHere,  we  adopted  this  general  framework  to  study learning  of direction  in  global \nmotion.  In contrast to Poggio et.  aI's supervised learning paradigm, we  used unsu(cid:173)\npervised learning both in the psychophysical experiments and modeling of learning. \nWe  designed a set of psychophysical tasks to study whether fast  learning occurs in \ndiscrimination of opposite directions of global  motion  and to explore the limits  of \nthis learning.  To  model the learning, we  studied two  models that differ in the way \nthey deal with noise. \n\n2  Psychophysics \n\nBall and Sekuler (1982, 1987) showed that discriminability of the direction of motion \nof two random dot patterns improved with training.  In this learning paradigm, more \nthan 2000 trials are required for reaching a stable performance.  Such a  \"slow\"  learn(cid:173)\ning time scale has been reported for  the learning of other perceptual tasks, such as \nvernier acuity (McKee and Westheimer 1978, Fahle 1994), stereoacuity (Fendick and \nWestheimer,  1983;  Ramachandran  and  Braddick  1973)  and  discrimination  of line \norientation (Vogels and Orban 1985).  In contrast to this  \"slow learning,\"  Fiorentini \nand  Berardi  (1981)  showed  that  for  learning  the  discrimination  of complex  grat(cid:173)\nings with two  harmonics of different spatial phase,  100-200  trials suffice.  Similarly \n(1992)  show  that  a  small  number  of trials  suffice  for  significantly \nPoggio  et.  al. \nimproving  performance on  a  vernier  hyperacuity task.  Both studies discussed  the \nspecificity of learning to the stimulus attributes. \nIn our study,  we  used a  two-alternative,  forced-choice  psychophysical procedure to \nmeasure the subject's ability to discriminate between two opposite directions of mo(cid:173)\ntion in dynamic random dot patterns in which 25% of the dots provide a correlated \n\n\fLearning  Direction  in  Global  Motion \n\n919 \n\nFigure  1:  (a)  Stimulus:  a  fraction  of the  dots  (filled  circles)  move  coherently  in  the \nsignal direction;  the rest (open circles) move randomly,  (b) Fast improvement is observed \non  the first  day  of testing,  and  retained  over  a  period  of time.  Each  block  consisted  of \n40  trials;  data  points  averaged  over 4 subjects,  and  errorbars  show standard  error. \n\nmotion signal spatially dispersed in a masking motion noise due to random motion \nof the reminder 75% of dots (Fig.  1(a)).  Each trial lasted 90 msec during which two \nframes were presented (with inter-frame interval equal to zero).  A session consisted \nof 4-6  blocks of 40  trials each.  Feedback was not provided during the experimental \nsessions.  Observers were required to maintain fixation  on a fixation  mark placed at \n2\u00b0  from  the imaginary circumference of the stimulus. \n\nTo  investigate the effects of practice and their retention, the discrimination of left(cid:173)\nward vs rightward direction of motion in the display was tested first on each exper(cid:173)\nimental session  for  3 consecutive days and  repeated  10  days  later.  The results  are \npresented in  Fig.  1(b).  For  most  observers,  a  fast  and  dramatic improvement  was \nseen  in  the first  day.  Fig.  1(b)  shows  that learning was  maintained in  subsequent \ndays,  and even ten days later without any training in  between.  Examination of the \nindividual observers' data revealed that improvement of performance occurred only \nif they started above chance level.  This suggests that this fast learning might imply \nthe improvement of an existing representation of the stimulus. \nIn  additional experiments,  we  did  not find  transfer to another direction  of motion \n(up/down),  indicating  that  the  learning  is  selective  to  specific  characteristics  of \nthe  stimulus.  Details  of experiments  testing  the  limits  of the  learning  appear  in \nVaina et. al  (1995). \n\n3  Modeling \n\nWe  propose  two  paradigms  to  model  the  learning  found  in  psychophysics.  Both \nuse  directionally-tuned  units  with  properties similar  to  those  of neurons  found  in \nMT (Maunsell and Van Essen, 1983).  Schematic MT neurons used in our modeling \nintegrate global  information by summing over  localized  responses: \n\n~ -(~(oi-od) \n\u2022 \n\nXt  =  ~e  2<1\" \n\ni=l \n\n(1) \n\nwhere  (Jh  is  the  standard  deviation  of the  tuning,  and  information  from  n  local \nresponses  is  taken  into  account.  The  responses  of a  collection  of such  units  (each \n\n\f920 \n\nV.  Sundareswaran,  Lucia M.  Vaina \n\n(a) \n\n(b) \n\nFigure 2:  Architectures  of the  models:  (a)  Learning  to Accommodate:  Cluster  Gaus(cid:173)\nsians  !I and  h  operate  on  input  vector  X  to  decide  global  motion  direction,  (b) \nLearning to  Ignore:  Global  motion  direction  is  computed  as  a weighted  combination  of \nunits'  preferred  directions. \n\ntuned to a  different direction)  form  the input vector to the models. \n\n3.1  Learning to accommodate \n\nThis model is based on grouping similar data.  In this paradigm, after learning,  the \nnetwork has an estimate of the contribution of noise,  and takes it into account when \nperforming  the  task  of discriminating  between  two  different  directions  of motion; \nso,  we say that the model  \"learns to accommodate\"  the noise. \nFig 2(a)  contains a  schematic of this model.  The representation vector consists  of \nresponses  Xl.X2, ..\u2022 , Xn  of the directionally-tuned units ..  Clustering is  done in the \nspace  of the representation vectors.  The model  is  a  combination of HyperBF-like \nfunctions  and  clustering  (Poggio  et.  aI,  1992,  Moody  and  Darken  1989).  We  use \ngaussians  with  mean  at  the  cluster  centers,  a.nd  \"move\"  the  cluster  centers  by  a \nlearning algorithm. \nA  cluster  gaussian  computes  a  gaussian  function  of the  representation  vector  for \nthe current stimulus from  the current center of the corresponding cluster.  We  say \n\"current\"  because the center is moved as the learning proceeds.  At any given point, \na  center  is  at  the  current  best  estimate  of the  center  of  the  corresponding  data \ncluster.  More precisely,  the learning rule to modify the jth coordinate of the center \nis  given  by: \n\nc(t-+:l)  =  c(t) . + '11 * (x~t) _  c(t) .) \n\n) \n\nW,) \n\n, \n\nW,) \n\nW,) \n\n'/ \n\n(2) \n\nwhere  w  is  the  index  of  \"winning\"  cluster  and  Cw,j  is  the  jth coordinate  of the \ncenter for  the wth cluster.  This moves  the center towards the new  data vector  X \nthat has  been judged  to  belong to the wth cluster.  The parameter TJ  controls the \nlearning rate. \n\n\fLearning  Direction  in  Global  Motion \n\n921 \n\n3.2  Learning to ignore \n\nThe learning  involved  in this approach  is  Hebbian,  and is  termed  \"learning to ig(cid:173)\nnore,\"  because  in  this weighted  averaging  scheme,  as  learning occurs,  the  weights \nfor  the noise response are progressively reduced to zero,  leaving only the contribu(cid:173)\ntion from  the signal.  In other words, the network learns to ignore the noise  (Vaina \net. al  1995). \nThe model's output is a global motion direction.  If the weights associated with the \nresponses  Xl, X2,\u00b7 . \u00b7 Xn  are WI, W2,'\"  W n ,  the global  motion direction is  calculated \nas \n\n-1 (L: wiXisinOi ) \n\n'\" \nw  WiXiCOSOj \n\n' \n\nII \n\nUo  =  tan \n\nwhere ti  is the tuned direction (angle) of the ith unit.  A schematic of the model is \nshown in Fig. 2(b).  The global direction is judged to be rightward if +Ot  > 80  > -Ot. \nWe  have examined  two  different  learning  rules:  exposure-based  learning,  and self(cid:173)\nsupervised learning. \nExposure-based learning:  The weight corresponding to a unit is incremented by \nan amount  proportional to  the current  weight.  Only  units whose  response  values \nare above a  certain threshold  are allowed  to  increase their weights.  This learning \nrule favors  units that fire  consistently: \n\n(3) \n\nwhere  rt  is a threshold,  and 1]  is  a small fraction  that controls the learning rate. \nSelf-supervised learning:  The weight corresponding to a  unit is increased by an \namount  proportional  to  the product  of the  current weight  and  a  decreasing  func(cid:173)\ntion  (exponential)  of the angular  difference  between  the calculated  global  motion \ndirection and the direction of tuning of the unit: \n\nWi  ~ Wi + 1]Wie (- (O; -OO)2/2u e) ,  if Xi> rt . \n\nIn this case,  the model  uses  its own estimate of the global  direction as  an  internal \nfeedback to determine the learning. \nAn  approach  similar  to  this model  for  learning vernier  hyperacuity  was  proposed \nby Weiss et.  al  (1993). \n\n4  Experiments \n\nFor the simulations, the input (motion) vectors were represented by their angles rel(cid:173)\native to the positive horizontal axis (Le., the magnitude is ignored).  The responses of \nthe directionally-tuned units can be directly computed from the angles (see Eqn.  1; \nalternatively, cosine tuning functions were used,  and similar results were obtained) . \nFor each trial, the coherent  motion  direction was  randomly decided.  In the exper(cid:173)\niments, O'h  was  set to  ~ ,  where  nd  is  the number of unit  preferred directions; the \ndirections  Ot  are chosen  by  uniformly dividing  nd  in to 211'.  In all the experiments, \neight  preferred  directions  were  used;  each  trial contained  40  random  dots  moving \nwith a correlation of25% (the same correlation as in the tests with human subjects); \n\nn d \n\n\f922 \n\nV.  Sundareswaran,  Lucia M.  Vaina \n\n1.00 \n\nI 0.90 \n\nu  0.80 \n8  0.70 \n'f \n~ 0.60 \n\nDo \n\n0.9 \n\n0.8 \n\n0.7 \n\n0.6 \n\n-G - Espoeure \n\n--.-- Self\u00b7npervised \n\nBlock No \n\n14 \n\nO.S \n\n0 \n\ns \n\n10 \n\n1S \n\n20 \n\nBlock No \n\nFigure 3:  Results from  typical  simulation  runs.  On  the left is  the curve for  the learning \nto  accommodate  model  (1}  = 0.005;  cluster  (j  = 0.5)  and  on  the  right  are  the  curves \nfor  the  learning  to  ignore  model. \n\neach block consisted of 50  trials.  The performance is  measured by fraction  correct, \ncorresponding to the fraction of the inputs that were correctly classified. \n\nFor  both  models,  we  did  simulations  to justify  the architecture of the  models  by \ndisabling learning,  and studying the performance of the models for  increasing cor(cid:173)\nrelation.  The performance was qUalitatively similar to that described in perceptual \nstudies of these stimuli in  humans and monkeys. \n\nLearning  curves  from  simulations  of both  models  are shown  in  Fig.  3.  These  are \nresults  from  averaging  over  the  performance  in  ten  simulation  runs.  As  can  be \nseen,  the  models  learn  to  do  the  discrimination.  While  both  models  successfully \nlearned to do the direction discrimination, there are quantitative differences  in  the \nlearning  curves.  The  learning  to  accommodate  paradigm  improved  very  rapidly \n(somewhat faster than the human subjects).  The exposure-based learning rule for the \nlearning to ignore paradigm learned at a rate comparable with the human observers. \nHowever, the final performance in this case was not very stable, and oscillated (not \nshown),  consistent with the observations of Weiss  et.  al  (1993)  in  learning vernier \nhyperacuity.  The  self-supervised  rule,  on  the  other  hand,  reached  a  stable  level \nof performance,  but the learning  was  slower;  the  reason  is  that  this  learning  rule \nexhibits instability if a  high value of learning rate (1})  is  used. \nIf either  model  trained  on  inputs  containing  horizontal  correlated  motion  was \nlater presented with inputs containing vertical correlated motion,  the performance \ndropped to pre-training levels.  This non-transfer  is consistent with  psychophysical \nresults mentioned in Section 2. \n\n5  Discussion \n\nIn this study, we  focused on learning of global motion direction, and not on motion \nperception.  Motion  perception  is  well-understood  both  physiologically  and  psy-\n\n\fLearning  Direction  in  Global Motion \n\n923 \n\nchophysically.  In a series of studies of neuronal correlates of the perceptual decision \nof direction  of motion  in  stochastic  motion  signals  like  the stimuli  used  here,  to(cid:173)\ngether with studies of performance on these stimuli on  monkeys with  MT lesions, \nNewsome and his collaborators (for a review,  Newsome et.  aI,  1989)  provide strong \nsupport for  the hypothesis that perceptual judgements of motion  direction  are  by \nand large based on  the directional signals carried by MT neurons. \nSalzman and  Newsome  (1994)  reported that in trained monkeys,  the perception of \nmotion in stochastic noise is  more  likely mediated by a winner-take-all mechanism \nthan by a weighted averaging mechanism.  Interestingly, our clustering model,  after \nlearning,  is  behaviorally similar to the winner-take-all mechanism:  the dominating \ncomponent of the representation results in the assignment to  the closest cluster. \nOur psychophysical  studies clearly demonstrate that perceptual learning of global \nmotion direction  occurs.  While  significant  progress  has  been  made  to understand \nthe  mechanisms  and  underlying  circuitry  of  learning  (Zohary  et.  al  1994)  as  yet \nthere is no satisfactory biologically explanation of how  and where this learning may \noccur  (Zohary and Newsome  1994). \nOur focus  in this paper was on computational models of learning direction in global \nmotion.  For  this,  we  proposed  two  approaches,  which  differ  in  the  way  they deal \nwith noise:  one learns to accommodates noise, and the other learns to ignore it.  We \ndo  not advocate that one or the other of the models we  proposed here provides the \nbiologically correct choice  for  this task.  However,  together with the psychophysics \ndescribed  here  these  models suggest  new  experiments which  we  are now  exploring \nboth psychophysically and computationally. \nWe  hope  that  by  closely  connecting  models  and  psychophysics  while  keeping  in \nmind  the  aim  of  neuronal  compatibility,  we  will  make  progress  in  understanding \nhow  the cortex learns so fast to discriminate direction of motion in extremely noisy \nsituations. \n\nAcknowledgments \n\nThis research was conducted at the Intelligent Systems Laboratory of Boston Uni(cid:173)\nversity,  College  of  Engineering.  LMV  and  VS  were  supported  in  part  by  grants \nfrom  the  Office  of Naval  Research  (# N00014-93-1-0381)  and  the  National  Insti(cid:173)\ntute of Health  (EY R01- 07861)  to LMV.  VS  was in  part supported by the Boston \nUniversity  College  of  Engineering  Dean's  postdoctoral  fellowship.  We  gratefully \nacknowledge  additional  financial  support  for  VS  from  the  Dean's  special  research \nfund. \n\nReferences \n\n[1]  K.  Ball  and  R.  Sekuler.  Direction-specific  improvement  in  motion  discrimination. \n\nVision  Research,  27(6):953-965,  1987. \n\n[2]  M.  Fahle.  Human  pattern  recognition:  parallel  processing  and perceptual  learning. \n\nPerception,  23:411-427,  1994. \n\n[3]  M.  Fendick and G.  Westheimer.  Effects of practice and the separation of test targets \n\non  foveal  and perifoveal hyperacuity.  Vision  Research,  23: 145-150,  1983. \n\n\f924 \n\nV.  Sundareswaran,  Lucia M.  Vaina \n\n[4]  A. Fiorentini and N.  Berardi.  Perceptual learning specific for  orientation and spatial \n\nfrequency.  Nature,  287(5777):43-44, September 1980. \n\n[5]  Y.  Fregnac,  D.  Shulz,  S.  Thorpe,  and  E.  Bienstock.  A  cellular  analogue  of visual \n\ncortical plasticity.  Nature,  333:367-370, 1988. \n\n[6]  E.  J.  Gibson. \n\nImprovements  in  perceptual  judgement  as  a  function  of controlled \n\npractice of training.  Psychology  bulletin, 50:402-431,  1953. \n\n[7]  C.  D.  Gilbert  and  T.  N.  Wiesel.  Receptive  field  dynamics  in  adult  primary  visual \n\ncortex.  Nature,  356:150-152,  1992. \n\n[8]  A.  Karni  and D.  Sagi.  Where practice makes perfect in texture discrimination:  Evi(cid:173)\ndence for  primary visual cortex plasticity.  Proc. Natl. Acad.  Sci.  USA,  88:4966- 4970, \nJune 1991. \n\n[9]  J. H. R.  Maunsell and D.  C.  Van Essen. Functional properties of neurons in the middle \ntemporal  visual  area  of the  macaque  monkey  i:  selectivity  for  stimulus  direction, \nspeed,  and orientation.  J.  Neurophysiology,  49:1127-1147,  1983. \n\n[10]  S.  P.  McKee  and G.  Westheimer.  Improvement in venier acuity with practice.  Per(cid:173)\n\nception  \u20ac3  Psychophysics,  24:258-262,  1978. \n\n[11]  W. T. Newsome,  K.  H.  Britten, and J.  A.  Movshon.  Neuronal correlates of a  percep(cid:173)\n\ntual decision.  Nature,  341 :52-54,  1989. \n\n[12]  T.  Poggio.  A theory of how  the brain might work.  In  Cold  Spring  Harbor  Symposia \non  Quantitative Biology,  pages 899-910.  Cold Spring Harbor Laboratory Press, 1990. \n[13]  T. Poggio,  M.  Fahle, and S.  Edelman.  Fast perceptual learning in visual hyperacuity. \n\nScience,  256:1018-1021,  1992. \n\n[14]  T.  Poggio  and F.  Girosi.  A theory of networks  for  approximation  and learning.  AI \n\nMemo 1140,  M.LT,  July 1989. \n\n[15]  V.  S.  Ramachandran  and  O.  Braddick.  Orientation-specific  learning  in  stereopsis. \n\nPerception,  2:371-376,  1973. \n\n[16]  C.  D.  Salzman  and  W.  T.  Newsome.  Neural  mechanisms  for  forming  a  perceptual \n\ndecision.  Science,  264:231-237,  1994. \n\n[17]  L.  M.  Vaina,  V.  Sundareswaran, and J.  Harris.  Computational learning and natural \n\nlearning.  Cognitive Brain Research,  1995.  (in press). \n\n[18]  L.  M. Vaina,  N.  M.  Grzywacz,  and M.  LeMay.  Structure from  motion with impaired \nlocal-speed  and  global  motion-field computations.  Neuml  Computation,  2:420-435, \n1990. \n\n[19]  R.  Vogels  and  G .  A.  Orban.  The  effect  of practice  on  the  oblique  effect  in  line \n\norientation judgements.  Vision  Research, 25:1679-1687,  1985. \n\n[20]  S.  Watamaniuk,  R.  Sekuler,  and D.  Williams.  Direction perception in complex  dy(cid:173)\nnamic displays:  the integration of direction information.  Vision  Research,  24:55- 62, \n1984. \n\n[21]  Y.  Weiss,  S.  Edelman, and M.  Fahle.  Models of perceptual learning in vernier hyper(cid:173)\n\nacuity.  Neural  Computation,  5:695-718,  1993. \n\n[22]  E. Zohary, S.  Celebrini, K. H.  Britten, and W. T . Newsome.  Neuronal plasticity that \nunderlies  improvement  in  perceptual  performance.  Science,  263:1289-1292,  March \n1994. \n\n[23]  E. Zohary and W. T. Newsome.  Perceptual learning in a direction discrimination task \n\nis  not based upon enhanced neuronal sensitivity in the sts.  Investigative  Ophthalmol(cid:173)\nogy  and  Visual  Science  supplement,  page 1663,  1994. \n\n\f", "award": [], "sourceid": 998, "authors": [{"given_name": "V.", "family_name": "Sundareswaran", "institution": null}, {"given_name": "Lucia", "family_name": "Vaina", "institution": null}]}