# Chef stability?

**URL:** <https://discourse.chef.io/t/chef-stability/1620>\
**Category:** Chef Infra (archive)\
**Created:** [November 17, 2010, 6:09pm UTC](https://discourse.chef.io/t/chef-stability/1620 "2010-11-17T18:09:04Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![Allan\_Carroll](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/allan_carroll/32/858_2.png) [@Allan\_Carroll](https://discourse.chef.io/u/Allan_Carroll)\
**Post date:** [November 17, 2010, 6:09pm UTC](https://discourse.chef.io/t/chef-stability/1620/1 "2010-11-17T18:09:04Z")

</div>

Hi,

I’ve been working the past few days on tweaking my chef scripts to go into  
production on EC2 and struggling to get anything I feel good about trusting.  
Chef looks like a great tool with a strong community. I’m hoping that there’s  
some Chef way of looking at the world I haven’t been exposed to that you can  
all enlighten me on.

I’m running Ubuntu 10.10 on EC2 with the version of chef from the Opscode Lucid  
repo (0.9.8).

A few things going on:

I can’t seem to keep chef-solr or chef-solr-indexer from crashing. I keep  
having to restart them for some reason. I’m using It makes everything feel  
really flakey, but I’m not convinced that’s the only thing I’m running into.

Sometimes the webui (and knife) show the status of all the nodes and sometimes  
it refuses saying that I have no nodes (even though the node list shows there  
are some there). The error in the logs is only the same 500 internal server  
error: connection refused that I see for lots of things.

Running chef-client by hand on a machine causes a different result than letting  
the timer driven version work. Like it forces the client to reevaluate all the  
data bags and search results and actually apply them.

Sometimes the clients get new data/nodes and update everything fine, sometimes  
they don’t.

Yesterday I started 8 boxes to bring a whole cluster up. On a few of them, Chef  
just randomly stopped working. Running chef-client by hand finished building  
the box correctly. One of them built part of a configuration file using data  
from a node that I had deleted off the Chef server a few hours earlier and then  
could never get out of that state. Deleting the configuration file and  
rerunning client fixed it.

Anyway, all of these small, but annoying, little glitches give me a really bad  
feeling about trusting Chef to manage my production infrastructure. Of the  
tools I’ve looked at, it’s the most promising.

I’d really like to given the promise of such powerful ability when it works,  
the time that I’ve put into it, and the time it will save. Is anyone using Chef  
at a large scale? Does it take handholding and massaging along the way, and  
that’s just the price for cutting-edge technology that will be solved as the  
code matures?

Thanks,  
Allan

---

<div class="post-metadata">

**Author:** ![Adam\_Jacob](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/adam_jacob/32/292_2.png) [@Adam\_Jacob](https://discourse.chef.io/u/Adam_Jacob)\
**Post date:** [November 17, 2010, 6:16pm UTC](https://discourse.chef.io/t/chef-stability/1620/2 "2010-11-17T18:16:19Z")

</div>

On Wed, Nov 17, 2010 at 10:09 AM, [allanca@gmail.com](mailto:allanca@gmail.com) wrote:

> I've been working the past few days on tweaking my chef scripts to go into  
> production on EC2 and struggling to get anything I feel good about trusting.  
> Chef looks like a great tool with a strong community. I'm hoping that there's  
> some Chef way of looking at the world I haven't been exposed to that you can  
> all enlighten me on.
> 
> I'm running Ubuntu 10.10 on EC2 with the version of chef from the Opscode Lucid  
> repo (0.9.8).
> 
> A few things going on:
> 
> I can't seem to keep chef-solr or chef-solr-indexer from crashing. I keep  
> having to restart them for some reason. I'm using It makes everything feel  
> really flakey, but I'm not convinced that's the only thing I'm running into.

This is going to be the source of several problems - can you send us a  
gist of what you get in the logs when these crash?

> Sometimes the webui (and knife) show the status of all the nodes and sometimes  
> it refuses saying that I have no nodes (even though the node list shows there  
> are some there). The error in the logs is only the same 500 internal server  
> error: connection refused that I see for lots of things.

Those pages both use search - if you are seeing consistent failures of  
Solr, thats the source of these issues.

> Running chef-client by hand on a machine causes a different result than letting  
> the timer driven version work. Like it forces the client to reevaluate all the  
> data bags and search results and actually apply them.

In what way? The code paths here are identical for the most part. If  
you're using data bags and search in the recipes, and you are seeing  
failures of Solr, I would wager that these differences are actually  
just a representation of the search service not being stable for you.

> Sometimes the clients get new data/nodes and update everything fine, sometimes  
> they don't.

Again, if it's data that comes from search, that's your issue.

> Yesterday I started 8 boxes to bring a whole cluster up. On a few of them, Chef  
> just randomly stopped working. Running chef-client by hand finished building  
> the box correctly. One of them built part of a configuration file using data  
> from a node that I had deleted off the Chef server a few hours earlier and then  
> could never get out of that state. Deleting the configuration file and  
> rerunning client fixed it.

All the symptoms you talk about sound search related - so we should  
focus there. 🙂

> Anyway, all of these small, but annoying, little glitches give me a really bad  
> feeling about trusting Chef to manage my production infrastructure. Of the  
> tools I've looked at, it's the most promising.

Sorry to hear that, but it's been quite stable for us (and for lots of  
other folks). We'll get you fixed up.

> I'd really like to given the promise of such powerful ability when it works,  
> the time that I've put into it, and the time it will save. Is anyone using Chef  
> at a large scale? Does it take handholding and massaging along the way, and  
> that's just the price for cutting-edge technology that will be solved as the  
> code matures?

There are people using Chef at the scale of many thousands of systems,  
and Opscode manages a production multi-tenant infrastructure that is  
also quite significant, using many of the same components that are in  
the open source Chef.

Happy to help - hook us up with the logs.

Best,  
Adam

--  
Opscode, Inc.  
Adam Jacob, CTO  
T: (206) 508-7449 E: [adam@opscode.com](mailto:adam@opscode.com)

---

<div class="post-metadata">

**Author:** ![Allan\_Carroll](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/allan_carroll/32/858_2.png) [@Allan\_Carroll](https://discourse.chef.io/u/Allan_Carroll)\
**Post date:** [November 17, 2010, 7:37pm UTC](https://discourse.chef.io/t/chef-stability/1620/3 "2010-11-17T19:37:52Z")

</div>

Whew. That makes it seem tractable. Thanks for helping zero in on this.

Here's what I dug up:

solr-indexer.log has no real clues.

Lots of these:

INFO: Indexing node 37192f37-447a-41c7-8480-c048c878743e from chef status error Connection refused - connect(2)}

and lots of these:

INFO: Indexing cookbook\_version 2bd0feeb-3e32-4bb2-867c-41e0cfa12806 from chef status ok}

solr.log also doesn't seem to have anything interesting, but here's the last set of output before it went away last time:

> <https://gist.github.com/allanca/703913>

Here's a typical failure from the server log:

> <https://gist.github.com/allanca/703912>

On Nov 17, 2010, at 11:16 AM, Adam Jacob wrote:

> On Wed, Nov 17, 2010 at 10:09 AM, [allanca@gmail.com](mailto:allanca@gmail.com) wrote:
> 
> > I've been working the past few days on tweaking my chef scripts to go into  
> > production on EC2 and struggling to get anything I feel good about trusting.  
> > Chef looks like a great tool with a strong community. I'm hoping that there's  
> > some Chef way of looking at the world I haven't been exposed to that you can  
> > all enlighten me on.
> > 
> > I'm running Ubuntu 10.10 on EC2 with the version of chef from the Opscode Lucid  
> > repo (0.9.8).
> > 
> > A few things going on:
> > 
> > I can't seem to keep chef-solr or chef-solr-indexer from crashing. I keep  
> > having to restart them for some reason. I'm using It makes everything feel  
> > really flakey, but I'm not convinced that's the only thing I'm running into.
> 
> This is going to be the source of several problems - can you send us a  
> gist of what you get in the logs when these crash?
> 
> > Sometimes the webui (and knife) show the status of all the nodes and sometimes  
> > it refuses saying that I have no nodes (even though the node list shows there  
> > are some there). The error in the logs is only the same 500 internal server  
> > error: connection refused that I see for lots of things.
> 
> Those pages both use search - if you are seeing consistent failures of  
> Solr, thats the source of these issues.
> 
> > Running chef-client by hand on a machine causes a different result than letting  
> > the timer driven version work. Like it forces the client to reevaluate all the  
> > data bags and search results and actually apply them.
> 
> In what way? The code paths here are identical for the most part. If  
> you're using data bags and search in the recipes, and you are seeing  
> failures of Solr, I would wager that these differences are actually  
> just a representation of the search service not being stable for you.
> 
> > Sometimes the clients get new data/nodes and update everything fine, sometimes  
> > they don't.
> 
> Again, if it's data that comes from search, that's your issue.
> 
> > Yesterday I started 8 boxes to bring a whole cluster up. On a few of them, Chef  
> > just randomly stopped working. Running chef-client by hand finished building  
> > the box correctly. One of them built part of a configuration file using data  
> > from a node that I had deleted off the Chef server a few hours earlier and then  
> > could never get out of that state. Deleting the configuration file and  
> > rerunning client fixed it.
> 
> All the symptoms you talk about sound search related - so we should  
> focus there. 🙂
> 
> > Anyway, all of these small, but annoying, little glitches give me a really bad  
> > feeling about trusting Chef to manage my production infrastructure. Of the  
> > tools I've looked at, it's the most promising.
> 
> Sorry to hear that, but it's been quite stable for us (and for lots of  
> other folks). We'll get you fixed up.
> 
> > I'd really like to given the promise of such powerful ability when it works,  
> > the time that I've put into it, and the time it will save. Is anyone using Chef  
> > at a large scale? Does it take handholding and massaging along the way, and  
> > that's just the price for cutting-edge technology that will be solved as the  
> > code matures?
> 
> There are people using Chef at the scale of many thousands of systems,  
> and Opscode manages a production multi-tenant infrastructure that is  
> also quite significant, using many of the same components that are in  
> the open source Chef.
> 
> Happy to help - hook us up with the logs.
> 
> Best,  
> Adam
> 
> --  
> Opscode, Inc.  
> Adam Jacob, CTO  
> T: (206) 508-7449 E: [adam@opscode.com](mailto:adam@opscode.com)

---

<div class="post-metadata">

**Author:** ![Blake\_Barnett](https://avatars.discourse-cdn.com/v4/letter/b/8e8cbc/32.png) [@Blake\_Barnett](https://discourse.chef.io/u/Blake_Barnett)\
**Post date:** [November 17, 2010, 11:52pm UTC](https://discourse.chef.io/t/chef-stability/1620/4 "2010-11-17T23:52:01Z")

</div>

I found that solr would crash reliably if the machine had a shortage of memory. If I increased the RAM allocated to the VM to ~2GB, it behaved much more reliably.

-Blake

On Nov 18, 2010, at 4:37 AM, Allan Carroll wrote:

> Whew. That makes it seem tractable. Thanks for helping zero in on this.
> 
> Here's what I dug up:
> 
> solr-indexer.log has no real clues.
> 
> Lots of these:
> 
> INFO: Indexing node 37192f37-447a-41c7-8480-c048c878743e from chef status error Connection refused - connect(2)}
> 
> and lots of these:
> 
> INFO: Indexing cookbook\_version 2bd0feeb-3e32-4bb2-867c-41e0cfa12806 from chef status ok}
> 
> solr.log also doesn't seem to have anything interesting, but here's the last set of output before it went away last time:
> 
> [solr.log output · GitHub](https://gist.github.com/703913)
> 
> Here's a typical failure from the server log:
> 
> [Typical server.log failure · GitHub](https://gist.github.com/703912)
> 
> On Nov 17, 2010, at 11:16 AM, Adam Jacob wrote:
> 
> > On Wed, Nov 17, 2010 at 10:09 AM, [allanca@gmail.com](mailto:allanca@gmail.com) wrote:
> > 
> > > I've been working the past few days on tweaking my chef scripts to go into  
> > > production on EC2 and struggling to get anything I feel good about trusting.  
> > > Chef looks like a great tool with a strong community. I'm hoping that there's  
> > > some Chef way of looking at the world I haven't been exposed to that you can  
> > > all enlighten me on.
> > > 
> > > I'm running Ubuntu 10.10 on EC2 with the version of chef from the Opscode Lucid  
> > > repo (0.9.8).
> > > 
> > > A few things going on:
> > > 
> > > I can't seem to keep chef-solr or chef-solr-indexer from crashing. I keep  
> > > having to restart them for some reason. I'm using It makes everything feel  
> > > really flakey, but I'm not convinced that's the only thing I'm running into.
> > 
> > This is going to be the source of several problems - can you send us a  
> > gist of what you get in the logs when these crash?
> > 
> > > Sometimes the webui (and knife) show the status of all the nodes and sometimes  
> > > it refuses saying that I have no nodes (even though the node list shows there  
> > > are some there). The error in the logs is only the same 500 internal server  
> > > error: connection refused that I see for lots of things.
> > 
> > Those pages both use search - if you are seeing consistent failures of  
> > Solr, thats the source of these issues.
> > 
> > > Running chef-client by hand on a machine causes a different result than letting  
> > > the timer driven version work. Like it forces the client to reevaluate all the  
> > > data bags and search results and actually apply them.
> > 
> > In what way? The code paths here are identical for the most part. If  
> > you're using data bags and search in the recipes, and you are seeing  
> > failures of Solr, I would wager that these differences are actually  
> > just a representation of the search service not being stable for you.
> > 
> > > Sometimes the clients get new data/nodes and update everything fine, sometimes  
> > > they don't.
> > 
> > Again, if it's data that comes from search, that's your issue.
> > 
> > > Yesterday I started 8 boxes to bring a whole cluster up. On a few of them, Chef  
> > > just randomly stopped working. Running chef-client by hand finished building  
> > > the box correctly. One of them built part of a configuration file using data  
> > > from a node that I had deleted off the Chef server a few hours earlier and then  
> > > could never get out of that state. Deleting the configuration file and  
> > > rerunning client fixed it.
> > 
> > All the symptoms you talk about sound search related - so we should  
> > focus there. 🙂
> > 
> > > Anyway, all of these small, but annoying, little glitches give me a really bad  
> > > feeling about trusting Chef to manage my production infrastructure. Of the  
> > > tools I've looked at, it's the most promising.
> > 
> > Sorry to hear that, but it's been quite stable for us (and for lots of  
> > other folks). We'll get you fixed up.
> > 
> > > I'd really like to given the promise of such powerful ability when it works,  
> > > the time that I've put into it, and the time it will save. Is anyone using Chef  
> > > at a large scale? Does it take handholding and massaging along the way, and  
> > > that's just the price for cutting-edge technology that will be solved as the  
> > > code matures?
> > 
> > There are people using Chef at the scale of many thousands of systems,  
> > and Opscode manages a production multi-tenant infrastructure that is  
> > also quite significant, using many of the same components that are in  
> > the open source Chef.
> > 
> > Happy to help - hook us up with the logs.
> > 
> > Best,  
> > Adam
> > 
> > --  
> > Opscode, Inc.  
> > Adam Jacob, CTO  
> > T: (206) 508-7449 E: [adam@opscode.com](mailto:adam@opscode.com)

---

<div class="post-metadata">

**Author:** ![Allan\_Carroll](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/allan_carroll/32/858_2.png) [@Allan\_Carroll](https://discourse.chef.io/u/Allan_Carroll)\
**Post date:** [November 18, 2010, 12:27am UTC](https://discourse.chef.io/t/chef-stability/1620/5 "2010-11-18T00:27:35Z")

</div>

That's likely the same problem I'm having. I've been trying to run my Chef server off of a machine with 700GB (EC2 micro instance).

This begs the larger question: what size of machine is recommended for running Chef? Seems like a pretty beefy system with all the parts running.

-Allan

On Nov 17, 2010, at 4:52 PM, Blake Barnett wrote:

> I found that solr would crash reliably if the machine had a shortage of memory. If I increased the RAM allocated to the VM to ~2GB, it behaved much more reliably.
> 
> -Blake
> 
> On Nov 18, 2010, at 4:37 AM, Allan Carroll wrote:
> 
> > Whew. That makes it seem tractable. Thanks for helping zero in on this.
> > 
> > Here's what I dug up:
> > 
> > solr-indexer.log has no real clues.
> > 
> > Lots of these:
> > 
> > INFO: Indexing node 37192f37-447a-41c7-8480-c048c878743e from chef status error Connection refused - connect(2)}
> > 
> > and lots of these:
> > 
> > INFO: Indexing cookbook\_version 2bd0feeb-3e32-4bb2-867c-41e0cfa12806 from chef status ok}
> > 
> > solr.log also doesn't seem to have anything interesting, but here's the last set of output before it went away last time:
> > 
> > [solr.log output · GitHub](https://gist.github.com/703913)
> > 
> > Here's a typical failure from the server log:
> > 
> > [Typical server.log failure · GitHub](https://gist.github.com/703912)
> > 
> > On Nov 17, 2010, at 11:16 AM, Adam Jacob wrote:
> > 
> > > On Wed, Nov 17, 2010 at 10:09 AM, [allanca@gmail.com](mailto:allanca@gmail.com) wrote:
> > > 
> > > > I've been working the past few days on tweaking my chef scripts to go into  
> > > > production on EC2 and struggling to get anything I feel good about trusting.  
> > > > Chef looks like a great tool with a strong community. I'm hoping that there's  
> > > > some Chef way of looking at the world I haven't been exposed to that you can  
> > > > all enlighten me on.
> > > > 
> > > > I'm running Ubuntu 10.10 on EC2 with the version of chef from the Opscode Lucid  
> > > > repo (0.9.8).
> > > > 
> > > > A few things going on:
> > > > 
> > > > I can't seem to keep chef-solr or chef-solr-indexer from crashing. I keep  
> > > > having to restart them for some reason. I'm using It makes everything feel  
> > > > really flakey, but I'm not convinced that's the only thing I'm running into.
> > > 
> > > This is going to be the source of several problems - can you send us a  
> > > gist of what you get in the logs when these crash?
> > > 
> > > > Sometimes the webui (and knife) show the status of all the nodes and sometimes  
> > > > it refuses saying that I have no nodes (even though the node list shows there  
> > > > are some there). The error in the logs is only the same 500 internal server  
> > > > error: connection refused that I see for lots of things.
> > > 
> > > Those pages both use search - if you are seeing consistent failures of  
> > > Solr, thats the source of these issues.
> > > 
> > > > Running chef-client by hand on a machine causes a different result than letting  
> > > > the timer driven version work. Like it forces the client to reevaluate all the  
> > > > data bags and search results and actually apply them.
> > > 
> > > In what way? The code paths here are identical for the most part. If  
> > > you're using data bags and search in the recipes, and you are seeing  
> > > failures of Solr, I would wager that these differences are actually  
> > > just a representation of the search service not being stable for you.
> > > 
> > > > Sometimes the clients get new data/nodes and update everything fine, sometimes  
> > > > they don't.
> > > 
> > > Again, if it's data that comes from search, that's your issue.
> > > 
> > > > Yesterday I started 8 boxes to bring a whole cluster up. On a few of them, Chef  
> > > > just randomly stopped working. Running chef-client by hand finished building  
> > > > the box correctly. One of them built part of a configuration file using data  
> > > > from a node that I had deleted off the Chef server a few hours earlier and then  
> > > > could never get out of that state. Deleting the configuration file and  
> > > > rerunning client fixed it.
> > > 
> > > All the symptoms you talk about sound search related - so we should  
> > > focus there. 🙂
> > > 
> > > > Anyway, all of these small, but annoying, little glitches give me a really bad  
> > > > feeling about trusting Chef to manage my production infrastructure. Of the  
> > > > tools I've looked at, it's the most promising.
> > > 
> > > Sorry to hear that, but it's been quite stable for us (and for lots of  
> > > other folks). We'll get you fixed up.
> > > 
> > > > I'd really like to given the promise of such powerful ability when it works,  
> > > > the time that I've put into it, and the time it will save. Is anyone using Chef  
> > > > at a large scale? Does it take handholding and massaging along the way, and  
> > > > that's just the price for cutting-edge technology that will be solved as the  
> > > > code matures?
> > > 
> > > There are people using Chef at the scale of many thousands of systems,  
> > > and Opscode manages a production multi-tenant infrastructure that is  
> > > also quite significant, using many of the same components that are in  
> > > the open source Chef.
> > > 
> > > Happy to help - hook us up with the logs.
> > > 
> > > Best,  
> > > Adam
> > > 
> > > --  
> > > Opscode, Inc.  
> > > Adam Jacob, CTO  
> > > T: (206) 508-7449 E: [adam@opscode.com](mailto:adam@opscode.com)

---

<div class="post-metadata">

**Author:** ![Adam\_Jacob](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/adam_jacob/32/292_2.png) [@Adam\_Jacob](https://discourse.chef.io/u/Adam_Jacob)\
**Post date:** [November 18, 2010, 12:55am UTC](https://discourse.chef.io/t/chef-stability/1620/6 "2010-11-18T00:55:19Z")

</div>

On Wed, Nov 17, 2010 at 4:27 PM, Allan Carroll [allanca@gmail.com](mailto:allanca@gmail.com) wrote:

> That's likely the same problem I'm having. I've been trying to run my Chef  
> server off of a machine with 700GB (EC2 micro instance).  
> This begs the larger question: what size of machine is recommended for  
> running Chef? Seems like a pretty beefy system with all the parts running.

Much of this will depend on what you are doing with it. Solr is going  
to want more ram as you add more indexed objects, and as the frequency  
with which you search increases. CouchDB tends to run quite leanly,  
and relies on OS file caching.

I've happily run a Chef server for nodes that check in every half on  
hour for around a 100 systems on an EC2 small instance.

Adam

--  
Opscode, Inc.  
Adam Jacob, CTO  
T: (206) 508-7449 E: [adam@opscode.com](mailto:adam@opscode.com)

---

<div class="post-metadata">

**Author:** ![Leinartas\_Michael](https://avatars.discourse-cdn.com/v4/letter/l/ac8455/32.png) [@Leinartas\_Michael](https://discourse.chef.io/u/Leinartas_Michael)\
**Post date:** [November 18, 2010, 1:03am UTC](https://discourse.chef.io/t/chef-stability/1620/7 "2010-11-18T01:03:57Z")

</div>

FWIW I was running chef-server 0.9.8 (and friends - rabbitmq, couchdb, solr) along with hosting a yum repo and an openvpn endpoint on a rackspace 512MB instance and having similar problems with chef-solr dying quite often once I reached 25 nodes or so. Updating to a 1GB instance solved it and I’m up to 50 nodes without trouble so far.

* * *

From: Allan Carroll [allanca@gmail.com](mailto:allanca@gmail.com)  
Reply-To: "chef@lists.opscode.com" [chef@lists.opscode.com](mailto:chef@lists.opscode.com)  
Date: Wed, 17 Nov 2010 18:27:35 -0600  
To: "chef@lists.opscode.com" [chef@lists.opscode.com](mailto:chef@lists.opscode.com)  
Subject: [chef] Re: Chef Server Hardware Reqs (was Re: Chef stability?)

That’s likely the same problem I’m having. I’ve been trying to run my Chef server off of a machine with 700GB (EC2 micro instance).

This begs the larger question: what size of machine is recommended for running Chef? Seems like a pretty beefy system with all the parts running.

-Allan

On Nov 17, 2010, at 4:52 PM, Blake Barnett wrote:

I found that solr would crash reliably if the machine had a shortage of memory. If I increased the RAM allocated to the VM to ~2GB, it behaved much more reliably.

-Blake

On Nov 18, 2010, at 4:37 AM, Allan Carroll wrote:

Whew. That makes it seem tractable. Thanks for helping zero in on this.

Here’s what I dug up:

solr-indexer.log has no real clues.

Lots of these:

INFO: Indexing node 37192f37-447a-41c7-8480-c048c878743e from chef status error Connection refused - connect(2)}

and lots of these:

INFO: Indexing cookbook\_version 2bd0feeb-3e32-4bb2-867c-41e0cfa12806 from chef status ok}

solr.log also doesn’t seem to have anything interesting, but here’s the last set of output before it went away last time:

> <https://gist.github.com/allanca/703913>

Here’s a typical failure from the server log:

> <https://gist.github.com/allanca/703912>

On Nov 17, 2010, at 11:16 AM, Adam Jacob wrote:

On Wed, Nov 17, 2010 at 10:09 AM, [allanca@gmail.com](mailto:allanca@gmail.com) wrote:  
I’ve been working the past few days on tweaking my chef scripts to go into  
production on EC2 and struggling to get anything I feel good about trusting.  
Chef looks like a great tool with a strong community. I’m hoping that there’s  
some Chef way of looking at the world I haven’t been exposed to that you can  
all enlighten me on.

I’m running Ubuntu 10.10 on EC2 with the version of chef from the Opscode Lucid  
repo (0.9.8).

A few things going on:

I can’t seem to keep chef-solr or chef-solr-indexer from crashing. I keep  
having to restart them for some reason. I’m using It makes everything feel  
really flakey, but I’m not convinced that’s the only thing I’m running into.

This is going to be the source of several problems - can you send us a  
gist of what you get in the logs when these crash?

Sometimes the webui (and knife) show the status of all the nodes and sometimes  
it refuses saying that I have no nodes (even though the node list shows there  
are some there). The error in the logs is only the same 500 internal server  
error: connection refused that I see for lots of things.

Those pages both use search - if you are seeing consistent failures of  
Solr, thats the source of these issues.

Running chef-client by hand on a machine causes a different result than letting  
the timer driven version work. Like it forces the client to reevaluate all the  
data bags and search results and actually apply them.

In what way? The code paths here are identical for the most part. If  
you’re using data bags and search in the recipes, and you are seeing  
failures of Solr, I would wager that these differences are actually  
just a representation of the search service not being stable for you.

Sometimes the clients get new data/nodes and update everything fine, sometimes  
they don’t.

Again, if it’s data that comes from search, that’s your issue.

Yesterday I started 8 boxes to bring a whole cluster up. On a few of them, Chef  
just randomly stopped working. Running chef-client by hand finished building  
the box correctly. One of them built part of a configuration file using data  
from a node that I had deleted off the Chef server a few hours earlier and then  
could never get out of that state. Deleting the configuration file and  
rerunning client fixed it.

All the symptoms you talk about sound search related - so we should  
focus there. 🙂

Anyway, all of these small, but annoying, little glitches give me a really bad  
feeling about trusting Chef to manage my production infrastructure. Of the  
tools I’ve looked at, it’s the most promising.

Sorry to hear that, but it’s been quite stable for us (and for lots of  
other folks). We’ll get you fixed up.

I’d really like to given the promise of such powerful ability when it works,  
the time that I’ve put into it, and the time it will save. Is anyone using Chef  
at a large scale? Does it take handholding and massaging along the way, and  
that’s just the price for cutting-edge technology that will be solved as the  
code matures?

There are people using Chef at the scale of many thousands of systems,  
and Opscode manages a production multi-tenant infrastructure that is  
also quite significant, using many of the same components that are in  
the open source Chef.

Happy to help - hook us up with the logs.

Best,  
Adam

–  
Opscode, Inc.  
Adam Jacob, CTO  
T: (206) 508-7449 E: [adam@opscode.com](mailto:adam@opscode.com)

---

<div class="post-metadata">

**Author:** ![Paul\_Paradise](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/paul_paradise/32/846_2.png) [@Paul\_Paradise](https://discourse.chef.io/u/Paul_Paradise)\
**Post date:** [November 18, 2010, 6:13am UTC](https://discourse.chef.io/t/chef-stability/1620/8 "2010-11-18T06:13:18Z")

</div>

On Wed, Nov 17, 2010 at 10:16 AM, Adam Jacob [adam@opscode.com](mailto:adam@opscode.com) wrote:

> On Wed, Nov 17, 2010 at 10:09 AM, [allanca@gmail.com](mailto:allanca@gmail.com) wrote:
> 
> > Running chef-client by hand on a machine causes a different result than  
> > letting  
> > the timer driven version work. Like it forces the client to reevaluate  
> > all the  
> > data bags and search results and actually apply them.
> 
> In what way? The code paths here are identical for the most part. If  
> you're using data bags and search in the recipes, and you are seeing  
> failures of Solr, I would wager that these differences are actually  
> just a representation of the search service not being stable for you.

This is almost undoubtedly unrelated to Allan's issues, but one repro-able  
example where chef runs differ when running in a daemon vs. single-shot is  
if your cookbook is complex and starts touching objects that outlive the  
lifespan of a single chef run -  
[http://tickets.opscode.com/browse/COOK-397for](http://tickets.opscode.com/browse/COOK-397for) example.

-Paul

---

<div class="post-metadata">

**Author:** ![Sean\_OMeara](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/sean_omeara/32/435_2.png) [@Sean\_OMeara](https://discourse.chef.io/u/Sean_OMeara)\
**Post date:** [November 18, 2010, 9:01am UTC](https://discourse.chef.io/t/chef-stability/1620/9 "2010-11-18T09:01:54Z")

</div>

thou shalt provide enough ram

-?

On Wed, Nov 17, 2010 at 8:03 PM, Leinartas, Michael  
[MICHAEL.LEINARTAS@orbitz.com](mailto:MICHAEL.LEINARTAS@orbitz.com) wrote:

> FWIW I was running chef-server 0.9.8 (and friends - rabbitmq, couchdb, solr)  
> along with hosting a yum repo and an openvpn endpoint on a rackspace 512MB  
> instance and having similar problems with chef-solr dying quite often once I  
> reached 25 nodes or so. Updating to a 1GB instance solved it and I'm up to  
> 50 nodes without trouble so far.
> 
> * * *
> 
> From: Allan Carroll [allanca@gmail.com](mailto:allanca@gmail.com)  
> Reply-To: "[chef@lists.opscode.com](mailto:chef@lists.opscode.com)" [chef@lists.opscode.com](mailto:chef@lists.opscode.com)  
> Date: Wed, 17 Nov 2010 18:27:35 -0600  
> To: "[chef@lists.opscode.com](mailto:chef@lists.opscode.com)" [chef@lists.opscode.com](mailto:chef@lists.opscode.com)  
> Subject: [chef] Re: Chef Server Hardware Reqs (was Re: Chef stability?)
> 
> That's likely the same problem I'm having. I've been trying to run my Chef  
> server off of a machine with 700GB (EC2 micro instance).
> 
> This begs the larger question: what size of machine is recommended for  
> running Chef? Seems like a pretty beefy system with all the parts running.
> 
> -Allan
> 
> On Nov 17, 2010, at 4:52 PM, Blake Barnett wrote:
> 
> I found that solr would crash reliably if the machine had a shortage of  
> memory. If I increased the RAM allocated to the VM to ~2GB, it behaved much  
> more reliably.
> 
> -Blake
> 
> On Nov 18, 2010, at 4:37 AM, Allan Carroll wrote:
> 
> Whew. That makes it seem tractable. Thanks for helping zero in on this.
> 
> Here's what I dug up:
> 
> solr-indexer.log has no real clues.
> 
> Lots of these:
> 
> INFO: Indexing node 37192f37-447a-41c7-8480-c048c878743e from chef status  
> error Connection refused - connect(2)}
> 
> and lots of these:
> 
> INFO: Indexing cookbook\_version 2bd0feeb-3e32-4bb2-867c-41e0cfa12806 from  
> chef status ok}
> 
> solr.log also doesn't seem to have anything interesting, but here's the last  
> set of output before it went away last time:
> 
> [solr.log output · GitHub](https://gist.github.com/703913)
> 
> Here's a typical failure from the server log:
> 
> [Typical server.log failure · GitHub](https://gist.github.com/703912)
> 
> On Nov 17, 2010, at 11:16 AM, Adam Jacob wrote:
> 
> On Wed, Nov 17, 2010 at 10:09 AM, [allanca@gmail.com](mailto:allanca@gmail.com) wrote:
> 
> I've been working the past few days on tweaking my chef scripts to go into  
> production on EC2 and struggling to get anything I feel good about trusting.  
> Chef looks like a great tool with a strong community. I'm hoping that  
> there's  
> some Chef way of looking at the world I haven't been exposed to that you can  
> all enlighten me on.
> 
> I'm running Ubuntu 10.10 on EC2 with the version of chef from the Opscode  
> Lucid  
> repo (0.9.8).
> 
> A few things going on:
> 
> I can't seem to keep chef-solr or chef-solr-indexer from crashing. I keep  
> having to restart them for some reason. I'm using It makes everything feel  
> really flakey, but I'm not convinced that's the only thing I'm running into.
> 
> This is going to be the source of several problems - can you send us a  
> gist of what you get in the logs when these crash?
> 
> Sometimes the webui (and knife) show the status of all the nodes and  
> sometimes  
> it refuses saying that I have no nodes (even though the node list shows  
> there  
> are some there). The error in the logs is only the same 500 internal server  
> error: connection refused that I see for lots of things.
> 
> Those pages both use search - if you are seeing consistent failures of  
> Solr, thats the source of these issues.
> 
> Running chef-client by hand on a machine causes a different result than  
> letting  
> the timer driven version work. Like it forces the client to reevaluate all  
> the  
> data bags and search results and actually apply them.
> 
> In what way? The code paths here are identical for the most part. If  
> you're using data bags and search in the recipes, and you are seeing  
> failures of Solr, I would wager that these differences are actually  
> just a representation of the search service not being stable for you.
> 
> Sometimes the clients get new data/nodes and update everything fine,  
> sometimes  
> they don't.
> 
> Again, if it's data that comes from search, that's your issue.
> 
> Yesterday I started 8 boxes to bring a whole cluster up. On a few of them,  
> Chef  
> just randomly stopped working. Running chef-client by hand finished building  
> the box correctly. One of them built part of a configuration file using data  
> from a node that I had deleted off the Chef server a few hours earlier and  
> then  
> could never get out of that state. Deleting the configuration file and  
> rerunning client fixed it.
> 
> All the symptoms you talk about sound search related - so we should  
> focus there. 🙂
> 
> Anyway, all of these small, but annoying, little glitches give me a really  
> bad  
> feeling about trusting Chef to manage my production infrastructure. Of the  
> tools I've looked at, it's the most promising.
> 
> Sorry to hear that, but it's been quite stable for us (and for lots of  
> other folks). We'll get you fixed up.
> 
> I'd really like to given the promise of such powerful ability when it works,  
> the time that I've put into it, and the time it will save. Is anyone using  
> Chef  
> at a large scale? Does it take handholding and massaging along the way, and  
> that's just the price for cutting-edge technology that will be solved as the  
> code matures?
> 
> There are people using Chef at the scale of many thousands of systems,  
> and Opscode manages a production multi-tenant infrastructure that is  
> also quite significant, using many of the same components that are in  
> the open source Chef.
> 
> Happy to help - hook us up with the logs.
> 
> Best,  
> Adam
> 
> --  
> Opscode, Inc.  
> Adam Jacob, CTO  
> T: (206) 508-7449 E: [adam@opscode.com](mailto:adam@opscode.com)

---

<div class="post-metadata">

**Author:** ![Chris\_Read](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/chris_read/32/859_2.png) [@Chris\_Read](https://discourse.chef.io/u/Chris_Read)\
**Post date:** [November 18, 2010, 9:30am UTC](https://discourse.chef.io/t/chef-stability/1620/10 "2010-11-18T09:30:38Z")

</div>

I was having solr crashes with 2GB RAM - turned out I needed to also  
increase the heap size to 512MB RAM to get things stable.

Chris

On Thu, Nov 18, 2010 at 1:03 AM, Leinartas, Michael \<  
[MICHAEL.LEINARTAS@orbitz.com](mailto:MICHAEL.LEINARTAS@orbitz.com)\> wrote:

> FWIW I was running chef-server 0.9.8 (and friends - rabbitmq, couchdb,  
> solr) along with hosting a yum repo and an openvpn endpoint on a rackspace  
> 512MB instance and having similar problems with chef-solr dying quite often  
> once I reached 25 nodes or so. Updating to a 1GB instance solved it and I'm  
> up to 50 nodes without trouble so far.
> 
> * * *
> 
> \*From: \*Allan Carroll [allanca@gmail.com](mailto:allanca@gmail.com)  
> \*Reply-To: \*"[chef@lists.opscode.com](mailto:chef@lists.opscode.com)" [chef@lists.opscode.com](mailto:chef@lists.opscode.com)  
> \*Date: \*Wed, 17 Nov 2010 18:27:35 -0600  
> \*To: \*"[chef@lists.opscode.com](mailto:chef@lists.opscode.com)" [chef@lists.opscode.com](mailto:chef@lists.opscode.com)  
> \*Subject: \*[chef] Re: Chef Server Hardware Reqs (was Re: Chef stability?)
> 
> That's likely the same problem I'm having. I've been trying to run my Chef  
> server off of a machine with 700GB (EC2 micro instance).
> 
> This begs the larger question: what size of machine is recommended for  
> running Chef? Seems like a pretty beefy system with all the parts running.
> 
> -Allan
> 
> On Nov 17, 2010, at 4:52 PM, Blake Barnett wrote:
> 
> I found that solr would crash reliably if the machine had a shortage of  
> memory. If I increased the RAM allocated to the VM to ~2GB, it behaved much  
> more reliably.
> 
> -Blake
> 
> On Nov 18, 2010, at 4:37 AM, Allan Carroll wrote:
> 
> Whew. That makes it seem tractable. Thanks for helping zero in on this.
> 
> Here's what I dug up:
> 
> solr-indexer.log has no real clues.
> 
> Lots of these:
> 
> INFO: Indexing node 37192f37-447a-41c7-8480-c048c878743e from chef status  
> error Connection refused - connect(2)}
> 
> and lots of these:
> 
> INFO: Indexing cookbook\_version 2bd0feeb-3e32-4bb2-867c-41e0cfa12806 from  
> chef status ok}
> 
> solr.log also doesn't seem to have anything interesting, but here's the  
> last set of output before it went away last time:
> 
> [solr.log output · GitHub](https://gist.github.com/703913)
> 
> Here's a typical failure from the server log:
> 
> [Typical server.log failure · GitHub](https://gist.github.com/703912)
> 
> On Nov 17, 2010, at 11:16 AM, Adam Jacob wrote:
> 
> On Wed, Nov 17, 2010 at 10:09 AM, [allanca@gmail.com](mailto:allanca@gmail.com) wrote:
> 
> I've been working the past few days on tweaking my chef scripts to go into  
> production on EC2 and struggling to get anything I feel good about  
> trusting.  
> Chef looks like a great tool with a strong community. I'm hoping that  
> there's  
> some Chef way of looking at the world I haven't been exposed to that you  
> can  
> all enlighten me on.
> 
> I'm running Ubuntu 10.10 on EC2 with the version of chef from the Opscode  
> Lucid  
> repo (0.9.8).
> 
> A few things going on:
> 
> I can't seem to keep chef-solr or chef-solr-indexer from crashing. I keep  
> having to restart them for some reason. I'm using It makes everything feel  
> really flakey, but I'm not convinced that's the only thing I'm running  
> into.
> 
> This is going to be the source of several problems - can you send us a  
> gist of what you get in the logs when these crash?
> 
> Sometimes the webui (and knife) show the status of all the nodes and  
> sometimes  
> it refuses saying that I have no nodes (even though the node list shows  
> there  
> are some there). The error in the logs is only the same 500 internal server  
> error: connection refused that I see for lots of things.
> 
> Those pages both use search - if you are seeing consistent failures of  
> Solr, thats the source of these issues.
> 
> Running chef-client by hand on a machine causes a different result than  
> letting  
> the timer driven version work. Like it forces the client to reevaluate all  
> the  
> data bags and search results and actually apply them.
> 
> In what way? The code paths here are identical for the most part. If  
> you're using data bags and search in the recipes, and you are seeing  
> failures of Solr, I would wager that these differences are actually  
> just a representation of the search service not being stable for you.
> 
> Sometimes the clients get new data/nodes and update everything fine,  
> sometimes  
> they don't.
> 
> Again, if it's data that comes from search, that's your issue.
> 
> Yesterday I started 8 boxes to bring a whole cluster up. On a few of them,  
> Chef  
> just randomly stopped working. Running chef-client by hand finished  
> building  
> the box correctly. One of them built part of a configuration file using  
> data  
> from a node that I had deleted off the Chef server a few hours earlier and  
> then  
> could never get out of that state. Deleting the configuration file and  
> rerunning client fixed it.
> 
> All the symptoms you talk about sound search related - so we should  
> focus there. 🙂
> 
> Anyway, all of these small, but annoying, little glitches give me a really  
> bad  
> feeling about trusting Chef to manage my production infrastructure. Of the  
> tools I've looked at, it's the most promising.
> 
> Sorry to hear that, but it's been quite stable for us (and for lots of  
> other folks). We'll get you fixed up.
> 
> I'd really like to given the promise of such powerful ability when it  
> works,  
> the time that I've put into it, and the time it will save. Is anyone using  
> Chef  
> at a large scale? Does it take handholding and massaging along the way, and  
> that's just the price for cutting-edge technology that will be solved as  
> the  
> code matures?
> 
> There are people using Chef at the scale of many thousands of systems,  
> and Opscode manages a production multi-tenant infrastructure that is  
> also quite significant, using many of the same components that are in  
> the open source Chef.
> 
> Happy to help - hook us up with the logs.
> 
> Best,  
> Adam
> 
> --  
> Opscode, Inc.  
> Adam Jacob, CTO  
> T: (206) 508-7449 E: [adam@opscode.com](mailto:adam@opscode.com)

---

<div class="post-metadata">

**Author:** ![Allan\_Carroll](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/allan_carroll/32/858_2.png) [@Allan\_Carroll](https://discourse.chef.io/u/Allan_Carroll)\
**Post date:** [November 18, 2010, 4:01pm UTC](https://discourse.chef.io/t/chef-stability/1620/11 "2010-11-18T16:01:20Z")

</div>

I moved to a larger instance and increased the heap size and have had much better results with solr too. Looks like it just needs more space.

On Nov 18, 2010, at 2:30 AM, Chris Read wrote:

> I was having solr crashes with 2GB RAM - turned out I needed to also increase the heap size to 512MB RAM to get things stable.
> 
> Chris
> 
> On Thu, Nov 18, 2010 at 1:03 AM, Leinartas, Michael [MICHAEL.LEINARTAS@orbitz.com](mailto:MICHAEL.LEINARTAS@orbitz.com) wrote:  
> FWIW I was running chef-server 0.9.8 (and friends - rabbitmq, couchdb, solr) along with hosting a yum repo and an openvpn endpoint on a rackspace 512MB instance and having similar problems with chef-solr dying quite often once I reached 25 nodes or so. Updating to a 1GB instance solved it and I'm up to 50 nodes without trouble so far.
> 
> From: Allan Carroll [allanca@gmail.com](mailto:allanca@gmail.com)  
> Reply-To: "[chef@lists.opscode.com](mailto:chef@lists.opscode.com)" [chef@lists.opscode.com](mailto:chef@lists.opscode.com)  
> Date: Wed, 17 Nov 2010 18:27:35 -0600  
> To: "[chef@lists.opscode.com](mailto:chef@lists.opscode.com)" [chef@lists.opscode.com](mailto:chef@lists.opscode.com)  
> Subject: [chef] Re: Chef Server Hardware Reqs (was Re: Chef stability?)
> 
> That's likely the same problem I'm having. I've been trying to run my Chef server off of a machine with 700GB (EC2 micro instance).
> 
> This begs the larger question: what size of machine is recommended for running Chef? Seems like a pretty beefy system with all the parts running.
> 
> -Allan
> 
> On Nov 17, 2010, at 4:52 PM, Blake Barnett wrote:
> 
> I found that solr would crash reliably if the machine had a shortage of memory. If I increased the RAM allocated to the VM to ~2GB, it behaved much more reliably.
> 
> -Blake
> 
> On Nov 18, 2010, at 4:37 AM, Allan Carroll wrote:
> 
> Whew. That makes it seem tractable. Thanks for helping zero in on this.
> 
> Here's what I dug up:
> 
> solr-indexer.log has no real clues.
> 
> Lots of these:
> 
> INFO: Indexing node 37192f37-447a-41c7-8480-c048c878743e from chef status error Connection refused - connect(2)}
> 
> and lots of these:
> 
> INFO: Indexing cookbook\_version 2bd0feeb-3e32-4bb2-867c-41e0cfa12806 from chef status ok}
> 
> solr.log also doesn't seem to have anything interesting, but here's the last set of output before it went away last time:
> 
> [solr.log output · GitHub](https://gist.github.com/703913)
> 
> Here's a typical failure from the server log:
> 
> [Typical server.log failure · GitHub](https://gist.github.com/703912)
> 
> On Nov 17, 2010, at 11:16 AM, Adam Jacob wrote:
> 
> On Wed, Nov 17, 2010 at 10:09 AM, [allanca@gmail.com](mailto:allanca@gmail.com) wrote:  
> I've been working the past few days on tweaking my chef scripts to go into  
> production on EC2 and struggling to get anything I feel good about trusting.  
> Chef looks like a great tool with a strong community. I'm hoping that there's  
> some Chef way of looking at the world I haven't been exposed to that you can  
> all enlighten me on.
> 
> I'm running Ubuntu 10.10 on EC2 with the version of chef from the Opscode Lucid  
> repo (0.9.8).
> 
> A few things going on:
> 
> I can't seem to keep chef-solr or chef-solr-indexer from crashing. I keep  
> having to restart them for some reason. I'm using It makes everything feel  
> really flakey, but I'm not convinced that's the only thing I'm running into.
> 
> This is going to be the source of several problems - can you send us a  
> gist of what you get in the logs when these crash?
> 
> Sometimes the webui (and knife) show the status of all the nodes and sometimes  
> it refuses saying that I have no nodes (even though the node list shows there  
> are some there). The error in the logs is only the same 500 internal server  
> error: connection refused that I see for lots of things.
> 
> Those pages both use search - if you are seeing consistent failures of  
> Solr, thats the source of these issues.
> 
> Running chef-client by hand on a machine causes a different result than letting  
> the timer driven version work. Like it forces the client to reevaluate all the  
> data bags and search results and actually apply them.
> 
> In what way? The code paths here are identical for the most part. If  
> you're using data bags and search in the recipes, and you are seeing  
> failures of Solr, I would wager that these differences are actually  
> just a representation of the search service not being stable for you.
> 
> Sometimes the clients get new data/nodes and update everything fine, sometimes  
> they don't.
> 
> Again, if it's data that comes from search, that's your issue.
> 
> Yesterday I started 8 boxes to bring a whole cluster up. On a few of them, Chef  
> just randomly stopped working. Running chef-client by hand finished building  
> the box correctly. One of them built part of a configuration file using data  
> from a node that I had deleted off the Chef server a few hours earlier and then  
> could never get out of that state. Deleting the configuration file and  
> rerunning client fixed it.
> 
> All the symptoms you talk about sound search related - so we should  
> focus there. 🙂
> 
> Anyway, all of these small, but annoying, little glitches give me a really bad  
> feeling about trusting Chef to manage my production infrastructure. Of the  
> tools I've looked at, it's the most promising.
> 
> Sorry to hear that, but it's been quite stable for us (and for lots of  
> other folks). We'll get you fixed up.
> 
> I'd really like to given the promise of such powerful ability when it works,  
> the time that I've put into it, and the time it will save. Is anyone using Chef  
> at a large scale? Does it take handholding and massaging along the way, and  
> that's just the price for cutting-edge technology that will be solved as the  
> code matures?
> 
> There are people using Chef at the scale of many thousands of systems,  
> and Opscode manages a production multi-tenant infrastructure that is  
> also quite significant, using many of the same components that are in  
> the open source Chef.
> 
> Happy to help - hook us up with the logs.
> 
> Best,  
> Adam
> 
> --  
> Opscode, Inc.  
> Adam Jacob, CTO  
> T: (206) 508-7449 E: [adam@opscode.com](mailto:adam@opscode.com)

---

<div class="post-metadata">

**Author:** ![Adam\_Jacob](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/adam_jacob/32/292_2.png) [@Adam\_Jacob](https://discourse.chef.io/u/Adam_Jacob)\
**Post date:** [November 18, 2010, 7:43pm UTC](https://discourse.chef.io/t/chef-stability/1620/12 "2010-11-18T19:43:04Z")

</div>

On Thu, Nov 18, 2010 at 1:30 AM, Chris Read [chris.read@gmail.com](mailto:chris.read@gmail.com) wrote:

> I was having solr crashes with 2GB RAM - turned out I needed to also  
> increase the heap size to 512MB RAM to get things stable.

Right - just having the RAM isn't enough, you need to tune the JVM as well.

For our (very large in comparison to almost everyone else, and way  
over-provisioned) setups config file:

solr\_java\_opts "-XX:MaxPermSize=1024m"  
solr\_heap\_size "6144M"  
solr\_java\_opts "-server"

That machine has 8GB of RAM.

Adam

--  
Opscode, Inc.  
Adam Jacob, CTO  
T: (206) 508-7449 E: [adam@opscode.com](mailto:adam@opscode.com)

---

<div class="post-metadata">

**Author:** ![Gilles\_Devaux](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/gilles_devaux/32/744_2.png) [@Gilles\_Devaux](https://discourse.chef.io/u/Gilles_Devaux)\
**Post date:** [November 18, 2010, 10:00pm UTC](https://discourse.chef.io/t/chef-stability/1620/13 "2010-11-18T22:00:50Z")

</div>

Maybe it's unrelated but if you don't install chef-server through the  
recipe you have to compact the couchdb indexes yourself or you'll  
waste tons of mem: [http://wiki.apache.org/couchdb/Compaction](http://wiki.apache.org/couchdb/Compaction)

Here is the script I use in a cron: [Compact chef's couchdb indexes · GitHub](https://gist.github.com/705726)

--Gilles

On Thu, Nov 18, 2010 at 11:43 AM, Adam Jacob [adam@opscode.com](mailto:adam@opscode.com) wrote:

> On Thu, Nov 18, 2010 at 1:30 AM, Chris Read [chris.read@gmail.com](mailto:chris.read@gmail.com) wrote:
> 
> > I was having solr crashes with 2GB RAM - turned out I needed to also  
> > increase the heap size to 512MB RAM to get things stable.
> 
> Right - just having the RAM isn't enough, you need to tune the JVM as well.
> 
> For our (very large in comparison to almost everyone else, and way  
> over-provisioned) setups config file:
> 
> solr\_java\_opts "-XX:MaxPermSize=1024m"  
> solr\_heap\_size "6144M"  
> solr\_java\_opts "-server"
> 
> That machine has 8GB of RAM.
> 
> Adam
> 
> --  
> Opscode, Inc.  
> Adam Jacob, CTO  
> T: (206) 508-7449 E: [adam@opscode.com](mailto:adam@opscode.com)
