# How to tell when recipes fail?

**URL:** <https://discourse.chef.io/t/how-to-tell-when-recipes-fail/1609>\
**Category:** Chef Infra (archive)\
**Created:** [November 8, 2010, 12:42am UTC](https://discourse.chef.io/t/how-to-tell-when-recipes-fail/1609 "2010-11-08T00:42:48Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mike\_Williams](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/mike_williams/32/854_2.png) [@Mike\_Williams](https://discourse.chef.io/u/Mike_Williams)\
**Post date:** [November 8, 2010, 12:42am UTC](https://discourse.chef.io/t/how-to-tell-when-recipes-fail/1609/1 "2010-11-08T00:42:48Z")

</div>

On the Chef web\_ui “Status” page, I can see when each node last checked-in with the server, but there does not appear to be any indication of failures during running of recipes!

What’s the recommended way of detecting and reporting on failures?

When chef-client uploads node state to the server, at the end of a run, does it persist any information about the success/failure of the run\_list?

Should I be registering an exception/report handler? And if so, has anyone written a handler that persists saves the run-reports to a central DB of some kind?

–  
cheers,  
Mike Williams

---

<div class="post-metadata">

**Author:** ![kallistec](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/kallistec/32/23_2.png) [@kallistec](https://discourse.chef.io/u/kallistec)\
**Post date:** [November 8, 2010, 4:49am UTC](https://discourse.chef.io/t/how-to-tell-when-recipes-fail/1609/2 "2010-11-08T04:49:26Z")

</div>

Ohai!

On Sun, Nov 7, 2010 at 4:42 PM, Mike Williams  
[mike@cogentconsulting.com.au](mailto:mike@cogentconsulting.com.au) wrote:

> Should I be registering an exception/report handler? And if so, has anyone written a handler that persists saves the run-reports to a central DB of some kind?

Yep, you want a report/exception handler. I started the wiki page for  
this feature earlier this week:  
[http://wiki.opscode.com/display/chef/Exception+and+Report+Handlers](http://wiki.opscode.com/display/chef/Exception+and+Report+Handlers)

Also, I've been working on a handler paired with a simple sinatra app  
that stores run history in redis. It's still pretty early going, but  
you can find the source here:  
[GitHub - danielsdeleo/nom\_nom\_nom: A Simple Status Server for Chef](https://github.com/danielsdeleo/nom_nom_nom) Just a word of warning,  
this hasn't been run in a production scenario anywhere, ever. I've  
been meaning to set it up for all of our non-production environments  
at opscode but I haven't had the time.

HTH,  
Dan DeLeo

> --  
> cheers,  
> Mike Williams

---

<div class="post-metadata">

**Author:** ![Mike\_Williams](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/mike_williams/32/854_2.png) [@Mike\_Williams](https://discourse.chef.io/u/Mike_Williams)\
**Post date:** [November 8, 2010, 5:25am UTC](https://discourse.chef.io/t/how-to-tell-when-recipes-fail/1609/3 "2010-11-08T05:25:02Z")

</div>

On 08/11/2010, at 15:49 , Daniel DeLeo wrote:

> Yep, you want a report/exception handler. I started the wiki page for  
> this feature earlier this week:  
> [http://wiki.opscode.com/display/chef/Exception+and+Report+Handlers](http://wiki.opscode.com/display/chef/Exception+and+Report+Handlers)
> 
> Also, I've been working on a handler paired with a simple sinatra app  
> that stores run history in redis. It's still pretty early going, but  
> you can find the source here:  
> [GitHub - danielsdeleo/nom\_nom\_nom: A Simple Status Server for Chef](https://github.com/danielsdeleo/nom_nom_nom)

Nice one, thanks Daniel.

Out of interest, why Redis, as opposed to CouchDB? I was thinking that it might make sense to store the result of the last Chef run in the existing server data-store, along with everything else, so that (for example) you could target failed nodes using "knife". Are you intentionally trying to keep "node status" and "run history" separate?

--  
cheers,  
Mike Williams

---

<div class="post-metadata">

**Author:** ![kallistec](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/kallistec/32/23_2.png) [@kallistec](https://discourse.chef.io/u/kallistec)\
**Post date:** [November 8, 2010, 6:36am UTC](https://discourse.chef.io/t/how-to-tell-when-recipes-fail/1609/4 "2010-11-08T06:36:09Z")

</div>

On Sun, Nov 7, 2010 at 9:25 PM, Mike Williams  
[mike@cogentconsulting.com.au](mailto:mike@cogentconsulting.com.au) wrote:

> On 08/11/2010, at 15:49 , Daniel DeLeo wrote:
> 
> > Also, I've been working on a handler paired with a simple sinatra app  
> > that stores run history in redis. It's still pretty early going, but  
> > you can find the source here:  
> > [GitHub - danielsdeleo/nom\_nom\_nom: A Simple Status Server for Chef](https://github.com/danielsdeleo/nom_nom_nom)
> 
> Nice one, thanks Daniel.
> 
> Out of interest, why Redis, as opposed to CouchDB? I was thinking that it might make sense to store the result of the last Chef run in the existing server data-store, along with everything else, so that (for example) you could target failed nodes using "knife". Are you intentionally trying to keep "node status" and "run history" separate?

I wrote it on the side as a way to get some exposure to technology I  
find interesting but don't use day-to-day. My plan is to get some  
experience using/operating what I have so far and then re-evaluate the  
technology decisions.

The app as it is allows you to fetch the list of only successful or  
failed nodes (it's part of how the data is modeled in redis) and the  
UI shows failed/success in the node list, so I'm not trying to keep  
them separate. But I did want to take the opportunity to design the UI  
(and the whole app, really) from first principles instead of trying to  
shoehorn it into what already exists so I can have a fresh perspective  
on the problem.

Cheers,

Dan DeLeo

> --  
> cheers,  
> Mike Williams

---

<div class="post-metadata">

**Author:** ![Rob\_Guttman](https://avatars.discourse-cdn.com/v4/letter/r/a9adbd/32.png) [@Rob\_Guttman](https://discourse.chef.io/u/Rob_Guttman)\
**Post date:** [November 8, 2010, 3:01pm UTC](https://discourse.chef.io/t/how-to-tell-when-recipes-fail/1609/5 "2010-11-08T15:01:21Z")

</div>

I've taken a different approach to monitoring failed recipes. I use a  
nagios passive check to monitor each chef-client's log file using LMF  
(similar to swatch) for both failed and successful runs. By adding a  
freshness threshold, it also doubles as a check that each chef-client is  
still running (i.e., the check goes critical if it hasn't received a passive  
check update within the chef client interval + splay + some buffer). Works  
well so far.

- Rob

On Mon, Nov 8, 2010 at 1:36 AM, Daniel DeLeo [dan@kallistec.com](mailto:dan@kallistec.com) wrote:

> On Sun, Nov 7, 2010 at 9:25 PM, Mike Williams  
> [mike@cogentconsulting.com.au](mailto:mike@cogentconsulting.com.au) wrote:
> 
> > On 08/11/2010, at 15:49 , Daniel DeLeo wrote:
> > 
> > > Also, I've been working on a handler paired with a simple sinatra app  
> > > that stores run history in redis. It's still pretty early going, but  
> > > you can find the source here:  
> > > [GitHub - danielsdeleo/nom\_nom\_nom: A Simple Status Server for Chef](https://github.com/danielsdeleo/nom_nom_nom)
> > 
> > Nice one, thanks Daniel.
> > 
> > Out of interest, why Redis, as opposed to CouchDB? I was thinking that  
> > it might make sense to store the result of the last Chef run in the existing  
> > server data-store, along with everything else, so that (for example) you  
> > could target failed nodes using "knife". Are you intentionally trying to  
> > keep "node status" and "run history" separate?
> 
> I wrote it on the side as a way to get some exposure to technology I  
> find interesting but don't use day-to-day. My plan is to get some  
> experience using/operating what I have so far and then re-evaluate the  
> technology decisions.
> 
> The app as it is allows you to fetch the list of only successful or  
> failed nodes (it's part of how the data is modeled in redis) and the  
> UI shows failed/success in the node list, so I'm not trying to keep  
> them separate. But I did want to take the opportunity to design the UI  
> (and the whole app, really) from first principles instead of trying to  
> shoehorn it into what already exists so I can have a fresh perspective  
> on the problem.
> 
> Cheers,
> 
> Dan DeLeo
> 
> > --  
> > cheers,  
> > Mike Williams

---

<div class="post-metadata">

**Author:** ![Sean\_OMeara](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/sean_omeara/32/435_2.png) [@Sean\_OMeara](https://discourse.chef.io/u/Sean_OMeara)\
**Post date:** [November 8, 2010, 3:15pm UTC](https://discourse.chef.io/t/how-to-tell-when-recipes-fail/1609/6 "2010-11-08T15:15:05Z")

</div>

I run chef-client from cron and check the exist status with $?

On Mon, Nov 8, 2010 at 10:01 AM, Rob Guttman [robguttman@gmail.com](mailto:robguttman@gmail.com) wrote:

> I've taken a different approach to monitoring failed recipes. I use a  
> nagios passive check to monitor each chef-client's log file using LMF  
> (similar to swatch) for both failed and successful runs. By adding a  
> freshness threshold, it also doubles as a check that each chef-client is  
> still running (i.e., the check goes critical if it hasn't received a passive  
> check update within the chef client interval + splay + some buffer). Works  
> well so far.
> 
> - Rob
> 
> On Mon, Nov 8, 2010 at 1:36 AM, Daniel DeLeo [dan@kallistec.com](mailto:dan@kallistec.com) wrote:
> 
> > On Sun, Nov 7, 2010 at 9:25 PM, Mike Williams  
> > [mike@cogentconsulting.com.au](mailto:mike@cogentconsulting.com.au) wrote:
> > 
> > > On 08/11/2010, at 15:49 , Daniel DeLeo wrote:
> > > 
> > > > Also, I've been working on a handler paired with a simple sinatra app  
> > > > that stores run history in redis. It's still pretty early going, but  
> > > > you can find the source here:  
> > > > [GitHub - danielsdeleo/nom\_nom\_nom: A Simple Status Server for Chef](https://github.com/danielsdeleo/nom_nom_nom)
> > > 
> > > Nice one, thanks Daniel.
> > > 
> > > Out of interest, why Redis, as opposed to CouchDB? I was thinking that  
> > > it might make sense to store the result of the last Chef run in the existing  
> > > server data-store, along with everything else, so that (for example) you  
> > > could target failed nodes using "knife". Are you intentionally trying to  
> > > keep "node status" and "run history" separate?
> > 
> > I wrote it on the side as a way to get some exposure to technology I  
> > find interesting but don't use day-to-day. My plan is to get some  
> > experience using/operating what I have so far and then re-evaluate the  
> > technology decisions.
> > 
> > The app as it is allows you to fetch the list of only successful or  
> > failed nodes (it's part of how the data is modeled in redis) and the  
> > UI shows failed/success in the node list, so I'm not trying to keep  
> > them separate. But I did want to take the opportunity to design the UI  
> > (and the whole app, really) from first principles instead of trying to  
> > shoehorn it into what already exists so I can have a fresh perspective  
> > on the problem.
> > 
> > Cheers,
> > 
> > Dan DeLeo
> > 
> > > --  
> > > cheers,  
> > > Mike Williams

---

<div class="post-metadata">

**Author:** ![Thibaut\_Barrere](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/thibaut_barrere/32/761_2.png) [@Thibaut\_Barrere](https://discourse.chef.io/u/Thibaut_Barrere)\
**Post date:** [November 8, 2010, 3:38pm UTC](https://discourse.chef.io/t/how-to-tell-when-recipes-fail/1609/7 "2010-11-08T15:38:22Z")

</div>

Hello!

fwiw, someone mentioned on IRC that they were planning to open-source a  
hoptoad exception handler that works with recent versions of chef.

I will definitely use that.

– Thibaut

---

<div class="post-metadata">

**Author:** ![Paul\_Choi](https://avatars.discourse-cdn.com/v4/letter/p/43a26b/32.png) [@Paul\_Choi](https://discourse.chef.io/u/Paul_Choi)\
**Post date:** [November 8, 2010, 10:26pm UTC](https://discourse.chef.io/t/how-to-tell-when-recipes-fail/1609/8 "2010-11-08T22:26:02Z")

</div>

I have a script that runs in cron, checking for the last time that a client has checked in.

I found that node[ohai\_time] records when a client has checked in. So if threshold \< Time.now - node[ohai\_time], notify.

This is the code that runs via cron on my servers:

> <https://gist.github.com/paulchoi/668386>

Please feel free to use it ifyou’d like, and feedback is welcome.

-Paul

-----Original Message-----  
From: Mike Williams [[mailto:mike@cogentconsulting.com.au](mailto:mike@cogentconsulting.com.au)]  
Sent: Sunday, November 07, 2010 4:43 PM  
To: [chef@lists.opscode.com](mailto:chef@lists.opscode.com)  
Subject: [chef] How to tell when recipes fail?

On the Chef web\_ui “Status” page, I can see when each node last checked-in with the server, but there does not appear to be any indication of failures during running of recipes!

What’s the recommended way of detecting and reporting on failures?

When chef-client uploads node state to the server, at the end of a run, does it persist any information about the success/failure of the run\_list?

Should I be registering an exception/report handler? And if so, has anyone written a handler that persists saves the run-reports to a central DB of some kind?

–  
cheers,  
Mike Williams

---

<div class="post-metadata">

**Author:** ![Mike\_Williams](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/mike_williams/32/854_2.png) [@Mike\_Williams](https://discourse.chef.io/u/Mike_Williams)\
**Post date:** [November 8, 2010, 11:00pm UTC](https://discourse.chef.io/t/how-to-tell-when-recipes-fail/1609/9 "2010-11-08T23:00:18Z")

</div>

On 09/11/2010, at 09:26 , Paul Choi wrote:

> I have a script that runs in cron, checking for the last time that a client has checked in.

That's definitely useful, but (correct me if I'm wrong) only ensures that the chef-client daemon is still running ... not that the recipes actually apply without error.

--  
cheers,  
Mike Williams

---

<div class="post-metadata">

**Author:** ![Paul\_Choi](https://avatars.discourse-cdn.com/v4/letter/p/43a26b/32.png) [@Paul\_Choi](https://discourse.chef.io/u/Paul_Choi)\
**Post date:** [November 9, 2010, 12:53am UTC](https://discourse.chef.io/t/how-to-tell-when-recipes-fail/1609/10 "2010-11-09T00:53:22Z")

</div>

Ah, you are right, sir... Monday brain fart on my part. 🙂 What I have been meaning to do is... write a reporter cookbook, which is run at the end of a cookbook, where it does an HTTP POST to Chef's data bags, and can be parsed for status.

Thanks for the feedback.

-Paul

-----Original Message-----  
From: Mike Williams [[mailto:mike@cogentconsulting.com.au](mailto:mike@cogentconsulting.com.au)]  
Sent: Monday, November 08, 2010 3:01 PM  
To: [chef@lists.opscode.com](mailto:chef@lists.opscode.com)  
Subject: [chef] Re: RE: How to tell when recipes fail?

On 09/11/2010, at 09:26 , Paul Choi wrote:

> I have a script that runs in cron, checking for the last time that a client has checked in.

That's definitely useful, but (correct me if I'm wrong) only ensures that the chef-client daemon is still running ... not that the recipes actually apply without error.

--  
cheers,  
Mike Williams

---

<div class="post-metadata">

**Author:** ![Gilles\_Devaux](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/gilles_devaux/32/744_2.png) [@Gilles\_Devaux](https://discourse.chef.io/u/Gilles_Devaux)\
**Post date:** [November 16, 2010, 5:40pm UTC](https://discourse.chef.io/t/how-to-tell-when-recipes-fail/1609/11 "2010-11-16T17:40:18Z")

</div>

Same here

On Mon, Nov 8, 2010 at 7:01 AM, Rob Guttman [robguttman@gmail.com](mailto:robguttman@gmail.com) wrote:

> I've taken a different approach to monitoring failed recipes. I use a  
> nagios passive check to monitor each chef-client's log file using LMF  
> (similar to swatch) for both failed and successful runs. By adding a  
> freshness threshold, it also doubles as a check that each chef-client is  
> still running (i.e., the check goes critical if it hasn't received a passive  
> check update within the chef client interval + splay + some buffer). Works  
> well so far.
> 
> - Rob
> 
> On Mon, Nov 8, 2010 at 1:36 AM, Daniel DeLeo [dan@kallistec.com](mailto:dan@kallistec.com) wrote:
> 
> > On Sun, Nov 7, 2010 at 9:25 PM, Mike Williams  
> > [mike@cogentconsulting.com.au](mailto:mike@cogentconsulting.com.au) wrote:
> > 
> > > On 08/11/2010, at 15:49 , Daniel DeLeo wrote:
> > > 
> > > > Also, I've been working on a handler paired with a simple sinatra app  
> > > > that stores run history in redis. It's still pretty early going, but  
> > > > you can find the source here:  
> > > > [GitHub - danielsdeleo/nom\_nom\_nom: A Simple Status Server for Chef](https://github.com/danielsdeleo/nom_nom_nom)
> > > 
> > > Nice one, thanks Daniel.
> > > 
> > > Out of interest, why Redis, as opposed to CouchDB? I was thinking that  
> > > it might make sense to store the result of the last Chef run in the existing  
> > > server data-store, along with everything else, so that (for example) you  
> > > could target failed nodes using "knife". Are you intentionally trying to  
> > > keep "node status" and "run history" separate?
> > 
> > I wrote it on the side as a way to get some exposure to technology I  
> > find interesting but don't use day-to-day. My plan is to get some  
> > experience using/operating what I have so far and then re-evaluate the  
> > technology decisions.
> > 
> > The app as it is allows you to fetch the list of only successful or  
> > failed nodes (it's part of how the data is modeled in redis) and the  
> > UI shows failed/success in the node list, so I'm not trying to keep  
> > them separate. But I did want to take the opportunity to design the UI  
> > (and the whole app, really) from first principles instead of trying to  
> > shoehorn it into what already exists so I can have a fresh perspective  
> > on the problem.
> > 
> > Cheers,
> > 
> > Dan DeLeo
> > 
> > > --  
> > > cheers,  
> > > Mike Williams
