Showing posts with label simulation. Show all posts
Showing posts with label simulation. Show all posts

Saturday, 9 December 2023

sequestering carbon, several books at a time CXXXV

 The latest, pre-Christmas, batch:


A third Checquy book!  I didn't know there was another one. Yay!

Thursday, 18 April 2019

Festschrift

400 pages of science: best birthday present ever!  Thanks so much, to everyone who contributed.  And special thanks to Andy and Viv, for pulling it all together.  Lots and lots of lovely reading ahead for me.  Yum.






Tuesday, 5 September 2017

ECAL 2017, Tuesday

Day two of the European Conference of Artificial Life, in Lyon, was the start of the conference proper.  The proceedings for the conference are open access.

We started with a fascinating keynote from philosopher Viola Schiaffonati, on Experimenting with computing and in computing: Stretching the traditional notion of experimentation.  As an ex-physicist, currently computer scientist, with a strong interest in computer simulation of biological and other complex systems, I found this a very useful exploration of the uses of simulation.  Physics-like subjects tend to have strong clear theories, and here “simulation” is more of “prediction without needing to run the experiment”, as in a virtual wind-tunnel, say.  Biology-like subjects, on the other hand, have less well-defined theories: they are more models comprising a “nest” of concepts, results and techniques from a range of sources, and simulations are more “does this nest hang together in correspondence with reality?”, or even “what tweaks do I need to make this nest hang together?”  The talk examined these concepts, put them together in an interesting way, and also had lots of useful references (not least to the “nest metaphor” paper which I need to read in detail).

After coffee I went to the Complex Dynamical Systems track.  Locating critical regions by the Relevance Index introduced a new measure useful for spotting systems at, or moving towards the “edge of chaos”.  Criticality as it could be demonstrated how agents who learn a structure that allows them to exhibit critical behaviour can exploit that learning to explore solutions to complex tasks in different environments.  Reservoir computing with a chaotic circuit was a neat demonstration of how even a (relatively) simple device can perform the complex computations of a reservoir computer.  Finally Signatures of criticality in a maximum entropy model of the C. elegans brain during free behaviour looked at potential critical behaviour in the “brain” of a (relatively!) simple biological organism.

smooth evolvable path surrounded by
rugged unevolvable peaks
After lunch I went to the Evolutionary Dynamics track.  Lineage selection leads to evolvability at large population sizes was a very nice talk demonstrating that the far ago ancestor of today’s creatures was not necessarily the fittest at the time: in a large population some groups can scamper up sharp fitness peaks and get stuck, whereas others trudge slowly up a smooth shallow incline of fitness, and survive to be even fitter in the long term. A 4-base model for the Aevol in-silico experimental evolution platform described the move from the classical binary representation to a four-base representation in the Aevol system, allowing degeneracy in the decoding.    MABE (Modular Agent Based Evolver): a framework for digital evolution research described a new modular architecture that should allow in silico evolution experiments to be set up much more straightforwardly.  Finally Gene duplications drive the evolution of complex traits and regulation described a series of experiments to investigate to various facets of gene duplication in Avida: is it the genetic material in order, or just the material, or just the increase in genome size, that’s important?

Then after another coffee break (hydration is very important at these events!) there was the second  keynote of the day: a fine double act from Andreas Wessel-Therhorn and Laurent Pujo-Menjouet speaking on The Illusion of Life, or the history and principles of animation.  They demon-strated the “12 principles of animation” with simple examples and then how they are realised in actual films.  They showed the clip of the death of Bambi’s mother that traumatised generations of children (myself included), showing that the actual death is never shown, despite the fact that many people remember it vividly. (I am not one of those who falsely recall it being depicted: I was traumatised by the fact of the death, not its depiction!) They also demonstrated how some of these animation principles have been used in the design of the Jibo robot’s movements, to give it “the illusion of life”.

The final event of the day was the poster session, including our Tuning Jordan algebra artificial chemistries with probability spawning functions.  I looked at all the posters quickly, but I then had to go off to the ISAL board meeting, where we reviewed the society’s activities, learned lessons from this year’s organisers, and brought next year’s organisers into the fold.





Thursday, 16 March 2017

not even pseudoscience

Sabine Hossenfelder nails it again – an argument against the simulation hypothesis from physics – but a much better one than usual.  The usual one tries to extrapolate physics from our universe to the “outside” one, which doesn’t work: they need not be the same.  Sabine argues about the physics of our universe within our universe: how hard it is to get consistent explanations, and why the hypothetical external programmer would have difficulties keeping up with our (simulated) scientists poking their noses into everything.

She’s a little grumpier than usual:
No, we probably don’t live in a computer simulation 
All this talk about how we might be living in a computer simulation pisses me off not because I’m afraid people will actually believe it. No, I think most people are much smarter than many self-declared intellectuals like to admit. Most readers will instead correctly conclude that today’s intelligencia is full of shit. And I can’t even blame them for it.



For all my social networking posts, see my Google+ page

Sunday, 16 October 2016

Elon Musk is wrong

Another argument against us living in a simulation:
Every day that you aren’t bombarded with styrofoam elephants or teleported to a planet made of cheese is another day that a potentially infinite number of software engineers have  decided not to play with the awesome thing that they built. 


For all my social networking posts, see my Google+ page

Monday, 12 September 2016

what's novelty?

Our paper “Defining and simulating open-ended novelty: requirements, guidelines, and challenges” has just been published in Theory in Biosciences.

abstract:
The open-endedness of a system is often defined as a continual production of novelty. Here we pin down this concept more fully by defining several types of novelty that a system may exhibit, classified as variation, innovation, and emergence. We then provide a meta-model for including levels of structure in a system’s model. From there, we define an architecture suitable for building simulations of open-ended novelty-generating systems and discuss how previously proposed systems fit into this framework. We discuss the design principles applicable to those systems and close with some challenges for the community.

A full-text view-only version of the paper is available from the publisher.


Monday, 11 July 2016

UCNC day 1

Another week, another conference.

I have moved from Cancun, Mexico, at the ALife conference, to Manchester, UK, for the conference on Unconventional Computation and Natural Computation (UCNC). It is very weird for me to be at a conference in the UK with “wrong way” jet lag!

The first day was a half day, starting with lunch – very civilised. The first talk was a tutorial from Jon Timmis on Swarm Robotics. This subject has multiple simple automomous robots working together with no global control, to produce an emergent behaviour and capability that none has individually. The tutorial covered the history of the subject, showing how some of the original constraints have become irrelevant: today’s “simple” robots are actually quite sophisticated compared to those at the discipline’s inception; and the original “nature inspiration” is no longer so prominent: use it if it helps, ignore it if it doesn’t. There are a couple of issues that make the subject difficult. The first is, how to design the local, individual robot rules that produce the desired emergent behaviour (and doesn’t produce undesired behaviours also)? This often reduces to an iterative design: suggest, test, refine, which can be automated in a search algorithm, such as an evolutionary search. This leads to the second issue: this search is most efficiently done in simulation, but there is a “reality gap” in simulation: the simulated physics is often too simplistic, leading to “overfitting” to the simulation and the solution then not working on the embodied physical robots. There are lots of fascinating results addressing these issues: the next challenge is moving this research out of the lab into the real world.

Then on to the technical session, with four talks. First up was my student, talking about using reservoir computing as an unconventional virtual machine for computing with carbon nanotubes: evolving the carbon nanotube system into a “good” reservoir, then training that reservoir to perform various tasks, rather than evolving the tasks directly. Next was another carbon nanotube talk, here having them in liquid crystal rather than frozen in polymer, allowing them to move to form clusters, to help their computational performance. The third talk changed tack, on to memristor logic. It appears that memristors natural support a ternary logic rather than the classical binary logic, and naturally implement different kinds of gates.  Finally, we had a talk that started with Zuse’s mechanical computer, and ended up with a “three cog, one gate” universal computer.

A great start to the conference.


Friday, 8 July 2016

ALife day 5

ALife day 5; last but not least.

The day started as usual with a fascinating keynote: today it was Linda Smith on “We need a developmental theory of environments”. Linda’s work is on development in human babies. She has gathered a rich corpus of information on babies’ perceived environments over their first two years of life. This has been gathered from head-mounted cameras (which today are so small they are just a chip in a headband), and demonstrate convincingly that the baby’s perceived environment changes dramatically over time, and that those changes are deeply embedded in its development. Early on, there are lots of close up faces, of a few adults. Later on, the baby’s view moves to hands: watching others, and its own, manipulating objects. Different experiments demonstrate the essential nature of the body / brain / environment feedback loop.  What is in this loop changes as the baby grows, and we need to understand when and how. And, of course (unless you are purely into how to experiment on babies for fun and profit), what does this tell us about developmental artificial life? The (perceived) environment is crucial to development.

I then went to the morning technical session on Artificial Chemistries, a potential substrate for ALife. We started with a talk on a novel replicator system based on a chemistry of functional combinators, with conservation of mass. The crucial design tradeoff is not to make the underlying artificial physics so strong that replication is trivial (a “copy organism” operation in the physics), nor to make it so sparse that replication is computationally infeasible. One way to strike the happy medium is to ensure the “functional units” are composed of a few “primitive units”, giving the system a small but crucial distance from the “atoms”. Next we heard about an extension of Hutton’s original replicator AChem, adding kinetics under the Gillespie algorithm, to find a “sweet spot” where a rich set of reaction occur in a computationally feasible time. Then we heard about “messy chemistries”, those that produce a wide range of uncontrolled products, and the conditions for one of the products to come to dominate, suggesting a “selection-first” AChem route to ALife. Then we had a description of a reaction-diffusion system incorporating energetics, and how a combination of exothermic and endothermic reaction systems can stabilise temperature across a region. Finally, we heard about taking mathematics seriously in order to use algebraic concepts, particularly non-associative algebras, to design a novel sub-symbolic AChem.

Then on to the closing keynote of the conference: Katie Bentley on “Do Endothelial Cells Dream of Eclectic Shape?” She explained the title: her work is about computational modelling of real biological systems, based on computational complex systems approaches. She had been warned biologists wouldn’t read something with the word “computational” in the title, so needed to use just biological words. But she wanted to signal to the CS-types that this might be of interest to them too, so used the punning title. She asked us if we got the pun: all but one hand went up. She then asked that person if they had seen the film; yes. She told us that if this was a straight biological conference, no one would have got the pun, and hardly anyone would have seen the film. Divided communities indeed. She went on to describe her computational model of vascular growth, in normal tissue and in tumours. Agent Based modelling, combined with real data and close interactions with biologists (who know which published results to trust, and which not), have resulted in several predictions that have been tested and confirmed in the wet lab. Mostly information flows from biology to ALife; this work demonstrates a great feedback from ALife into biology.

Then it was all over bar the closing ceremony: information about the International Society for ALife, the next two ALife conferences (ECAL 2017 in Lyon, France; ALife 2018 in Japan), and a variety of awards for best papers, lifetime achievements, and contributions to the community.

A truly excellent conference, in content and in organisation. I had a wonderful time, and my head is buzzing with ideas and connections. My neural pathways have been exercised and reconfigured. I need to go home and process all this information further.

Next year in Lyon.


Thursday, 7 July 2016

ALife day 4

ALife day 4. I’m at that point in conference-going where I keep having to check what day it is, as I’ve lost track.  Apparently it’s Thursday.  I never could get the hang of Thursdays.

This particular Thursday started brilliantly. Keynote speaker Randall Beer talked about “Autopoiesis and Enaction in the Game of Life”. Autopoiesis, or system self-production and self-maintenance, can be a slippery concept to explain. Beer takes the Game of Life, and uses it as a “toy” system, to get a handle on the concepts. The key is to look at the GoL from a process perspective (a toy “chemistry”), rather than from the more usual automaton perspective (the lower-level GoL toy “physics”). (This new modelling perspective is emergent under our definition, as it is actually a new meta-model.) In this different view, a glider can be seen as a very primitive autopoietic system, a network of processes constituting a well-defined identity that is self-maintained. In a marvellous tour de force, Beer showed how all the components fit together, and how GoL can be used to explain and illuminate concepts from autopoiesis in a beautifully clear and elegant manner.

The morning technical session was on Development (although not all the papers were). We started off with a discussion of the relationship between developmental encodings and hierarchical modular structure. Next; simulation is crucial for many ALife experiments, and MecaCell is one being built for developmental experiments. Then our paper; Bioreflective architectures generalise and combine the concepts of computational reflection and von Neumann’s Universal Constructor. Finally; it can be difficult to evolve heterogeneous specialist cooperative behaviours unless the specialists can recognise their partners.

The afternoon keynote was Francisco Santos talking about “Climate Change Governance, Cooperation and Self-organization”. This was an application of game theory in large finite populations to cooperation and coordination problems. He showed some counter-intuitive results: it is easier to get global cooperation via small groups initially, and it is better for small groups to invoke actions than to leave enforcement to a global organisation. In the end, the results show the old adage: think globally, act locally”, and that there is yet hope for cooperation.

I then went to the technical session on “Living technology and Human-Computer interaction”. First there were a couple of talks on the EvoBot, a modular liquid handling robot, covering both the hardware and the software. This system, being built as part of the EvoBliss EU project, is an open source design that costs two orders of magnitude less than current laboratory systems. It allows programmable chemical experiments, with precise, repetitive, complex operations. Real-time droplet identification allows complex operations to be specified. For example, it can be programmed to apply a droplet, wait for the droplet to start moving with a certain speed, then suck the droplet back up again; or to apply a chemical once an array of droplets has clustered. Next we learned about developing a computational agent to “play” a cooperative herding game with novice humans: some of the humans learned how to play from the agent, some never did, and most thought they were playing with another person. Finally we heard about a NetLogo-based approach to teaching school children about complex systems principles.

My brain is full, and there is still a day to go!

Wednesday, 3 February 2016

simulation the CoSMoS way

not this cover, sadly
We are currently writing a book about how to design, build, and use computer simulations “as scientific instruments”, using out CoSMoS (Complex Systems Modelling and Simulation) approach.  This is taking somewhat longer than planned, as other things always seem to have higher priority.

One of those other things is the CellBranch research project, funded by BBSRC, which finished last year.  We’ve just published a technical report describing in detail what we’ve done on the simulation side.

We used the CoSMoS approach, as described in that book in preparation.  The structured approach helped us enormously, particularly in getting to grips with the biological model we were simulating.  The report describes three increments of the simulation, and how to access the software developed.

One process step we invented, which is not part of the official approach in the current draft of the book, was a thing we dubbed the to don’t list.  We built the system incrementally, starting with a very simple version, then systematically adding the needed complexity.  While we were developing each increment, we kept having ideas about what would be needed next.  We wanted to remember these, but we also wanted to make it crystal clear that they were not to be included in the current increment.  So they got added to the to don’t list.

This turned out not to be just a helpful aide-memoire, but had an interesting side effect.  One feature of the CoSMoS approach is listing and justifying the assumptions being made during modelling and development.  Such assumptions are always made, but are often not documented, and so are promptly forgotten.  Making assumptions explicit, and teasing out their consequences, helps communication within a multidisciplinary team (here biologists and software engineers): “oh, that means we won’t be able to do X”.  And it helps enormously in subsequent increments, if one of the earlier, no longer forgotten, assumptions becomes invalidated because of new assumptions.  Nearly everything that went onto the to don’t list could be cast as an assumption: Y on the to don’t list could become “Y isn’t needed this increment, because…”.  This gave us much better insight into what the current increment could actually provide in terms of scientific understanding.

So now the to don’t list is being added to the official approach.  Maybe it’s a good job we hadn’t delivered the book on the original schedule, else we wouldn’t be able to make this valuable addition!  I hope the acquiring editor sees it the same way…

Wednesday, 10 June 2015

How many times should you run your simulation?

You have a new algorithm, say an evolutionary algorithm.  You want to show it is better than another algorithm.  But its behaviour is stochastic, varying from run to run.  How many times do you need to run it to show what you want to show?  How big a sample of possible results do you need? 10?  100?  1000?  More?

This is a known problem, solved by statisticians, used by scientists who want to know how many experiments they need to run to investigate their hypotheses.

First I recall some statistical terminology (null hypothesis, statistical significance, statistical power, effect size), then describe the way to calculate how many runs you need, using this terminology.  If you know this terminology, you can skip ahead to the final table.

null hypothesis

The null hypothesis, H0, is usually a statement of the status quo, of the null effect: the treatment has no effect; the algorithm has the same performance; etc.  The alternative hypothesis, H1, is that the effect is not null.  You are (usually) seeking to refute the null hypothesis in favour of the alternative hypothesis.

statistical significance and statistical power

Given a null hypothesis H0 and a statistical test, there are four possibilities.  H0 may or may not be true, and the test may or may not refute it.

H0 true H0 false
refute H0 type I error, false positive correct
fail to refute H0 correct type II error, false negative

The type I error, or false positive, erroneously concludes that the null hypothesis is false; that is, that there is an effect when there is not. The probability of this error (of refuting H0 given that H0 is true) is called α, or the statistical significance level.

     Î± = prob(refute H0|H0) = prob(false positive)

The type II error, or false negative, erroneously fails to refute the null hypothesis; that is, it is a failure to detect an effect that exists.  The probability of this error (of not refuting H0, given that H0 is not true) is called β.

     Î² = prob(not refute H0|not H0) = prob(false negative)

1−β is called the statistical power of the test, the probability of correctly refuting H0 when H0 is not true.

     power = 1−β = prob(refute H0|not H0)

Clearly, we would like to minimise both Î± and β, minimising both false positives and false negatives.  However, these aims are in conflict.  For example, we could trivially minimise Î± by never refuting H0, but this would maximise β, and vice versa.  The smaller either of them needs to be, the more runs are needed.

Different problems might put different emphasis on Î± and β.  Consider the case of illness diagnosis, versus your new algorithm.

Diagnosic test for an illness

  • false positive: the test detect illness when there is none, which might result in worry and unnecessary treatment
  • false negative: the test fails to detect the illness, which might result in death
So minimising Î² may take preference in such a case.

New evolutionary algorithm

  • false positive: you claim your algorithm is better when it isn’t; this will be embarrassing for you later when more runs fail to support your claim
  • false negative: you fail to detect that your algorithm is better; so you don’t publish
So minimising Î± may take preference in such a case.

p-value

A statistical test results in a p-value.  p is the probability of observing an effect, given that the null hypothesis holds.

     p = prob(obs|H0)

A low p-value means a low probability of observing the effect if H0 holds; the conclusion is then that H0 (probably) doesn’t hold.  We refute H0 with confidence level 1−p.

To refute H0 we want the observed p-value to be less than the initially chosen statistical significance Î±.  A typical confidence level is 95%, or Î± = 0.05.

effect size

The more runs we make, the smaller Î± and Î² can be.  That means we can refute H0 at higher and higher confidence level.  In fact, with enough runs we can refute almost any null hypothesis at a given confidence level, since any change to an algorithm probably makes some difference to the outcome.

Alternatively, we can use the higher number of runs to detect smaller and smaller deviations from H0.

For example, if H0 is that the means of the two samples A and B are the same:

     H0 : μA = μB

and the alternative hypothesis H1 is that the means are different:

     H1 : μA ≠ μB

then with more runs we can detect smaller and smaller differences in the means at a given significance level.  We can measure a smaller and smaller effect size.

Cohen’s d is one measure of effect size for normally distributed samples: the difference in means normalised by the standard deviation:

     d = (μA − μB) / σ

So an effect size of 1 occurs when the means are separated by one standard deviation.  Cohen gives names to different effect sizes: he calls 0.2 “small”, 0.5 “medium”, and 0.8 “large” effect sizes.  The smaller the effect size you want to detect, the more runs you need.

d = 0.2, “small”

d = 0.5, “medium”

d = 0.8, “large”

“Small” effect sizes might be fine in certain cases, where a small difference can nevertheless be valuable (either in a product, or in a theory).  However, for an algorithm to be worth publishing, you should probably aim for at least a “medium” effect size, if not a “large” one.

sample size

The required sample size, or number of runs, n, depends on your desired values of three parameters: (i) statistical significance Î±, (ii) statistical power 1−β, and (iii) effect size d.

For two samples of the same size, normally distributed, with H0 : Î¼A = Î¼B, then

     n = 2( (z(1−α/2)+z(1−β)) / d)2

where z is the inverse normal cumulative probability distribution function.  In the case of a single sample A compared against a fixed mean value, H0 : Î¼A = Î¼0, this number can be halved.

A Python script that calculates this value of n (rounded up to an integer) is
from math import ceil
from scipy.stats import norm
n = ceil( 2 * ( (norm.ppf ppf(1-alpha/2) + norm.ppf ppf(power) ) / d ) **2 )


If you want a significance at the 95% level (α = 0.05), and a power of 90%, and want to detect a large effect, then you need 33 runs; if you want to detect a medium effect you need 85 runs, and to detect a small effect you need over 500 runs.  (These numbers are based on the assumption that your samples are normally distributed.  Similar results follow for other distributions.)

At first sight, it might seem strange that the number of runs is somehow independent of the “noisiness” of the algorithm: why don’t we need more runs for an inherently noisy algorithm (large standard deviation) than from a very stable one (small standard deviation)?  The reason is because the effect size is also dependent on the standard deviation: for a given difference in the means, the effect size decreases as the standard deviation increases.  

So there you have it.  You don’t have to guess the number of runs, or just do “lots” to be sure: you can calculate n.  And it’s probably less than you would have guessed.



Wednesday, 29 April 2015

simulating manifestos

Want to check out the consequences of all the different parties’ election manifestos?  Run a simulation.

tl;dr: they all end in disaster.  Thank heavens governments never actually keep their manifesto pledges, in that case!


[via BoingBoing]

For all my social networking posts, see my Google+ page